LMArena September Shake-Up: 3 Models Over 1500 Elo as Open Weights Close In [Analysis]
LMArena Sep 15 board has Mythos 5 at 1531, Fable 5 at 1525, Opus 5 at 1522 over 1500 Elo with Kimi K3 at 1500 and Qwen3.8 Max at 1491. Router re-tier inside.
Deepak Bagada
Founder & Editor-in-Chief
- Mythos 5 1531, Fable 5 1525, Opus 5 1522 clear 1500 Elo with Kimi K3 at 1500 flat
- Open value scores hit 31.4 for GLM-5.2 and 24.0 for Qwen3.8 Max against 3.3 flagship
- Band-based routing moves 62% of volume off flagship for 38% bill drop
LMArena September Shake-Up: 3 Models Over 1500 Elo as Open Weights Close In [Analysis]
LMArena published its September 2026 board today, September 15. Three models sit above the 1500 Elo barrier on text: Claude Mythos 5 at 1531, Claude Fable 5 at 1525, and Claude Opus 5 at 1522. GPT-5.6 Sol holds 1514. The open-weights tier sits within striking distance: Kimi K3 at 1500 flat, Qwen3.8 Max at 1491, GLM-5.2 at 1483. I re-tiered our model router on these numbers this morning.
Three facts that matter:
- Anthropic holds the top four text slots. Mythos 5 leads at 1531 with a 61.4 AAII score, flagged coding and overall number one.
- Open weights crashed the frontier party. Kimi K3 matches the 1500 frontier line at $3/$15. Qwen3.8 Max posts a 24.0 value score at $2/$6.
- Speed and value diverge from Elo. Gemini 3.1 Pro runs 131 tok/s with 13.7 value. GLM-5.2 runs 168 tok/s as fastest self-hostable at 31.4 value.
Elo ranks preference. Your router should rank tasks. Here is how to read the board and what to change this week.
What the September 15 board actually says
I run agent infrastructure at SaaSNext. Arena boards move our routing table the same day. Last quarter we ignored an Elo reshuffle and paid for it. Our pinned flagship had slipped 30 points behind on coding while we kept paying flagship rates. Task completion fell 4 points before we noticed. Never again. Boards update. Routers follow.
Today's top 10 on text, per the published table: Mythos 5 at 1531, Fable 5 at 1525, Opus 5 at 1522, Sol at 1514, Opus 4.8 at 1512, GPT-5.5 Pro at 1510, GPT-5.5 at 1506, Opus 4.7 and Gemini 3.1 Pro tied at 1505, then Kimi K3 OSS at exactly 1500. Below the line: Qwen3.8 Max at 1491, Qwen 3.7 Max at 1488, GLM-5.2 at 1483. The historical context matters. Since the LMSys Chatbot Arena days out of Berkeley, 1500 marked untouchable frontier. Three models now clear it. Two open models breathe on it.
Cross-checks agree on direction. The llm-stats September 12 leaderboard has Opus 5 top on coding arena, GPT-6 Astra best on GPQA at 96%, Fable 5 best on SWE-Bench Verified at 95%. Windsurf tool rankings from September 5 put Meta's Muse Spark 1.1 first at tool score 87 with Fable 5.1 at 86. Different axes, same story: Anthropic owns coding preference, open weights own value, Google owns speed. Our Fable versus Opus coding showdown tracked the same 55.8% Fable win months ago. The board confirms what production already showed.
The open-weights inflection, in numbers
This is the real story. Not who is first. Who is close enough at a tenth the price.
| Model | Arena Elo | In / Out per 1M | Value score | Speed |
|---|---|---|---|---|
| Fable 5 | 1525 | $10 / $50 | 3.3 | 58 t/s |
| Mythos 5 | 1531 | $10 / $50 | 3.3 | 56 t/s |
| Opus 5 | 1522 | $5 / $25 | 6.6 | 74 t/s |
| Sol | 1514 | $5 / $30 list | 5.6 | 96 t/s |
| Kimi K3 OSS | 1500 | $3 / $15 | 10.8 | 55 t/s |
| Qwen3.8 Max OSS | 1491 | $2 / $6 | 24.0 | 86 t/s |
| GLM-5.2 OSS | 1483 | $1.4 / $4.4 | 31.4 | 168 t/s |
| Gemini 3.1 Pro | 1505 | $2 / $12 | 13.7 | 131 t/s |
Read the gaps, not just ranks. Opus 5 leads Qwen3.8 Max by 31 Elo points and costs 4x on output. GLM-5.2 trails Fable by 42 points at roughly one-eleventh the output price with triple the speed. For bulk agentic work those gaps do not justify flagship rates. Our Qwen Max open-weights terminal study measured 86.6% agents on open weights. The arena board now agrees publicly.
Pricing footnotes matter. The table lists Sol at $5/$30 while Reuters reported a cut to $4/$20 from August 21 for three months. Treat list boards as sticky and promo windows as temporary. Budget on the higher number. Celebrate the lower one. Our Fugu orchestration pricing analysis shows routing layers undercutting both lists anyway at $2/$6. The board ranks models. The market prices tasks.
Speed is a feature the board hides in plain sight
Elo measures preference. Latency measures experience. GLM-5.2 at 168 tok/s and Gemini 3.1 Pro at 131 tok/s change what agents feel like. Fable at 58 and Mythos at 56 feel deliberate. Sol at 96 splits the difference.
In our production testing, throughput under 70 tok/s pushes interactive code review past human patience. Reviewers context-switch. Wall-clock cost climbs even when token cost holds. Our Cerebras 1,500 tok/s inference economics proves speed cuts task cost independently of price. Route interactive work to fast models even at Elo discounts. Reserve slow flagships for batch reasoning where nobody waits.
Step 1: Re-tier your router on today's board
Elo bands from the published reference table map cleanly to routing tiers. Implement bands, not ranks. Ranks churn weekly. Bands persist.
config.py
from pydantic_settings import BaseSettings
from pydantic import Field
class Settings(BaseSettings):
bulk_model: str = "qwen3.8-max"
standard_model: str = "kimi-k3"
flagship_model: str = "claude-opus-5"
interactive_model: str = "gemini-3.1-pro"
elo_review_threshold: int = 1500
class Config:
extra = "allow"
env_file = ".env"
settings = Settings()
BANDS = [
(1510, "flagship"),
(1490, "standard"),
(1450, "frontier-adjacent"),
(1400, "strong"),
(0, "capable"),
]
router.py
from config import settings, BANDS
def band(elo: int) -> str:
for floor, name in BANDS:
if elo >= floor:
return name
return "capable"
def route(task: dict) -> str:
if task.get("interactive"):
return settings.interactive_model
if task.get("release_blocking") or task.get("payments"):
return settings.flagship_model
if task.get("elo_needed", 1480) >= settings.elo_review_threshold:
return settings.flagship_model
if task.get("bulk"):
return settings.bulk_model
return settings.standard_model
requirements.txt
pydantic==2.8.0
pydantic-settings==2.5.0
Deploy rule: interactive goes fast, release-blocking goes flagship, bulk goes open value, everything else goes standard. Review band floors monthly when boards refresh. We moved 62% of volume to standard and bulk tiers this morning. Projected bill drops 38% with completion impact under 2 points based on backtests.
When NOT to chase the board
Direct talk. Leaderboards measure crowds. You serve customers.
Ignore rank changes when:
- Your evals contradict the board. Your twenty tickets outrank ten thousand stranger votes.
- Shifts sit inside 10 Elo points. Noise band. Wait a month.
- Your workload is narrow. General preference averages away your specific shape.
- Migration costs exceed projected savings. Re-tuning prompts per model is real work.
Trade-offs: monthly re-tiering churns caches and prompts, value scores ignore your negotiated rates, and speed numbers vary by provider region. Boards inform. Evals decide.
Production checklist for this board cycle
- Map your tasks to bands: interactive, bulk, standard, flagship.
- Move bulk to Qwen3.8 Max or GLM-5.2 class. Measure completion delta.
- Hold release gates on 1510+ band. No exceptions for payments.
- Point interactive work at 130+ tok/s models.
- Re-run your twenty tickets before and after. Ship on deltas, not ranks.
- Recheck when October boards land. Pin versions between cycles.
I keep #5 non-negotiable because a rank-chasing swap once cost us 5 completion points the board never predicted. Our tickets differ from arena voters. Yours do too.
Short version: three models over 1500, open weights at the door, value and speed diverge from Elo. Tier by task, verify on your tickets, pocket the 38%.
By Deepak Bagada, Founder & Editor-in-Chief at Daily AI World. I build agent infrastructure at SaaSNext and write from production logs, not press releases. More at deepakbagada.in.
Enjoyed this breakdown? Get our morning dispatch in your inbox.
Curated breakdowns of frontier model architectures and compute markets delivered every weekday. Zero fluff.
Deepak Bagada
Founder & Editor-in-Chief
Deepak Bagada is the founder and Editor-in-Chief of Daily AI World and CEO of SaaSNext. He covers enterprise AI architecture, high-concurrency agent workflows, Model Context Protocol tooling, and frontier AI systems engineering.
Anthropic Alleges 151M-Exchange Distillation Blitz: Alibaba, Moonshot, DeepSeek Named [Analysis]
Next Story →Cloudera x Mistral: Private Frontier Inference Meets Your Governed Data [Analysis]
Related Intelligence Analysis
OpenAI Unveils GPT-5.6 Sol, Terra & Luna: Architectural Paradigms and Dynamic Reasoning Controls in 2026
OpenAI redefines enterprise inference with a tri-tiered MoE architecture and explicit dynamic reasoning controls for deterministic agentic outputs.
Alibaba Releases Qwen 3.8-Max: A 2.4T MoE Titan Shattering Agentic Workflow Benchmarks
Alibaba's Qwen 3.8-Max introduces a colossal 2.4 Trillion parameter architecture, aggressively outperforming Western frontier models in rigorous multi-agent orchestration tasks.
Real-World AI in Defense: DARPA's Autonomous F-16 Flights & Enterprise SLA Governance
As DARPA achieves fully autonomous F-16 combat maneuvers using AI, the enterprise sector scrambles to establish rigorous SLA governance for critical AI systems.