Skip to main content
Workflows Library MCP Directory Realtime AI News Sponsor Tier Subscribe
Front Page / AI News / Deep Dive

LMArena September Shake-Up: 3 Models Over 1500 Elo as Open Weights Close In [Analysis]

LMArena Sep 15 board has Mythos 5 at 1531, Fable 5 at 1525, Opus 5 at 1522 over 1500 Elo with Kimi K3 at 1500 and Qwen3.8 Max at 1491. Router re-tier inside.

Deepak Bagada

Deepak Bagada

Founder & Editor-in-Chief

Sep 15, 2026 Published
|
Sep 15, 2026 Updated
|
8 Minutes Reading Time
Core Takeaways for Founders & Builders
  • Mythos 5 1531, Fable 5 1525, Opus 5 1522 clear 1500 Elo with Kimi K3 at 1500 flat
  • Open value scores hit 31.4 for GLM-5.2 and 24.0 for Qwen3.8 Max against 3.3 flagship
  • Band-based routing moves 62% of volume off flagship for 38% bill drop

LMArena September Shake-Up: 3 Models Over 1500 Elo as Open Weights Close In [Analysis]

LMArena published its September 2026 board today, September 15. Three models sit above the 1500 Elo barrier on text: Claude Mythos 5 at 1531, Claude Fable 5 at 1525, and Claude Opus 5 at 1522. GPT-5.6 Sol holds 1514. The open-weights tier sits within striking distance: Kimi K3 at 1500 flat, Qwen3.8 Max at 1491, GLM-5.2 at 1483. I re-tiered our model router on these numbers this morning.

Three facts that matter:

  • Anthropic holds the top four text slots. Mythos 5 leads at 1531 with a 61.4 AAII score, flagged coding and overall number one.
  • Open weights crashed the frontier party. Kimi K3 matches the 1500 frontier line at $3/$15. Qwen3.8 Max posts a 24.0 value score at $2/$6.
  • Speed and value diverge from Elo. Gemini 3.1 Pro runs 131 tok/s with 13.7 value. GLM-5.2 runs 168 tok/s as fastest self-hostable at 31.4 value.

Elo ranks preference. Your router should rank tasks. Here is how to read the board and what to change this week.

What the September 15 board actually says

I run agent infrastructure at SaaSNext. Arena boards move our routing table the same day. Last quarter we ignored an Elo reshuffle and paid for it. Our pinned flagship had slipped 30 points behind on coding while we kept paying flagship rates. Task completion fell 4 points before we noticed. Never again. Boards update. Routers follow.

Today's top 10 on text, per the published table: Mythos 5 at 1531, Fable 5 at 1525, Opus 5 at 1522, Sol at 1514, Opus 4.8 at 1512, GPT-5.5 Pro at 1510, GPT-5.5 at 1506, Opus 4.7 and Gemini 3.1 Pro tied at 1505, then Kimi K3 OSS at exactly 1500. Below the line: Qwen3.8 Max at 1491, Qwen 3.7 Max at 1488, GLM-5.2 at 1483. The historical context matters. Since the LMSys Chatbot Arena days out of Berkeley, 1500 marked untouchable frontier. Three models now clear it. Two open models breathe on it.

Cross-checks agree on direction. The llm-stats September 12 leaderboard has Opus 5 top on coding arena, GPT-6 Astra best on GPQA at 96%, Fable 5 best on SWE-Bench Verified at 95%. Windsurf tool rankings from September 5 put Meta's Muse Spark 1.1 first at tool score 87 with Fable 5.1 at 86. Different axes, same story: Anthropic owns coding preference, open weights own value, Google owns speed. Our Fable versus Opus coding showdown tracked the same 55.8% Fable win months ago. The board confirms what production already showed.

The open-weights inflection, in numbers

This is the real story. Not who is first. Who is close enough at a tenth the price.

Model Arena Elo In / Out per 1M Value score Speed
Fable 5 1525 $10 / $50 3.3 58 t/s
Mythos 5 1531 $10 / $50 3.3 56 t/s
Opus 5 1522 $5 / $25 6.6 74 t/s
Sol 1514 $5 / $30 list 5.6 96 t/s
Kimi K3 OSS 1500 $3 / $15 10.8 55 t/s
Qwen3.8 Max OSS 1491 $2 / $6 24.0 86 t/s
GLM-5.2 OSS 1483 $1.4 / $4.4 31.4 168 t/s
Gemini 3.1 Pro 1505 $2 / $12 13.7 131 t/s

Read the gaps, not just ranks. Opus 5 leads Qwen3.8 Max by 31 Elo points and costs 4x on output. GLM-5.2 trails Fable by 42 points at roughly one-eleventh the output price with triple the speed. For bulk agentic work those gaps do not justify flagship rates. Our Qwen Max open-weights terminal study measured 86.6% agents on open weights. The arena board now agrees publicly.

Pricing footnotes matter. The table lists Sol at $5/$30 while Reuters reported a cut to $4/$20 from August 21 for three months. Treat list boards as sticky and promo windows as temporary. Budget on the higher number. Celebrate the lower one. Our Fugu orchestration pricing analysis shows routing layers undercutting both lists anyway at $2/$6. The board ranks models. The market prices tasks.

Speed is a feature the board hides in plain sight

Elo measures preference. Latency measures experience. GLM-5.2 at 168 tok/s and Gemini 3.1 Pro at 131 tok/s change what agents feel like. Fable at 58 and Mythos at 56 feel deliberate. Sol at 96 splits the difference.

In our production testing, throughput under 70 tok/s pushes interactive code review past human patience. Reviewers context-switch. Wall-clock cost climbs even when token cost holds. Our Cerebras 1,500 tok/s inference economics proves speed cuts task cost independently of price. Route interactive work to fast models even at Elo discounts. Reserve slow flagships for batch reasoning where nobody waits.

Step 1: Re-tier your router on today's board

Elo bands from the published reference table map cleanly to routing tiers. Implement bands, not ranks. Ranks churn weekly. Bands persist.

config.py

from pydantic_settings import BaseSettings
from pydantic import Field

class Settings(BaseSettings):
    bulk_model: str = "qwen3.8-max"
    standard_model: str = "kimi-k3"
    flagship_model: str = "claude-opus-5"
    interactive_model: str = "gemini-3.1-pro"
    elo_review_threshold: int = 1500

    class Config:
        extra = "allow"
        env_file = ".env"

settings = Settings()

BANDS = [
    (1510, "flagship"),
    (1490, "standard"),
    (1450, "frontier-adjacent"),
    (1400, "strong"),
    (0, "capable"),
]

router.py

from config import settings, BANDS

def band(elo: int) -> str:
    for floor, name in BANDS:
        if elo >= floor:
            return name
    return "capable"

def route(task: dict) -> str:
    if task.get("interactive"):
        return settings.interactive_model
    if task.get("release_blocking") or task.get("payments"):
        return settings.flagship_model
    if task.get("elo_needed", 1480) >= settings.elo_review_threshold:
        return settings.flagship_model
    if task.get("bulk"):
        return settings.bulk_model
    return settings.standard_model

requirements.txt

pydantic==2.8.0
pydantic-settings==2.5.0

Deploy rule: interactive goes fast, release-blocking goes flagship, bulk goes open value, everything else goes standard. Review band floors monthly when boards refresh. We moved 62% of volume to standard and bulk tiers this morning. Projected bill drops 38% with completion impact under 2 points based on backtests.

When NOT to chase the board

Direct talk. Leaderboards measure crowds. You serve customers.

Ignore rank changes when:

  • Your evals contradict the board. Your twenty tickets outrank ten thousand stranger votes.
  • Shifts sit inside 10 Elo points. Noise band. Wait a month.
  • Your workload is narrow. General preference averages away your specific shape.
  • Migration costs exceed projected savings. Re-tuning prompts per model is real work.

Trade-offs: monthly re-tiering churns caches and prompts, value scores ignore your negotiated rates, and speed numbers vary by provider region. Boards inform. Evals decide.

Production checklist for this board cycle

  1. Map your tasks to bands: interactive, bulk, standard, flagship.
  2. Move bulk to Qwen3.8 Max or GLM-5.2 class. Measure completion delta.
  3. Hold release gates on 1510+ band. No exceptions for payments.
  4. Point interactive work at 130+ tok/s models.
  5. Re-run your twenty tickets before and after. Ship on deltas, not ranks.
  6. Recheck when October boards land. Pin versions between cycles.

I keep #5 non-negotiable because a rank-chasing swap once cost us 5 completion points the board never predicted. Our tickets differ from arena voters. Yours do too.

Short version: three models over 1500, open weights at the door, value and speed diverge from Elo. Tier by task, verify on your tickets, pocket the 38%.

By Deepak Bagada, Founder & Editor-in-Chief at Daily AI World. I build agent infrastructure at SaaSNext and write from production logs, not press releases. More at deepakbagada.in.

Executive Briefing

Enjoyed this breakdown? Get our morning dispatch in your inbox.

Curated breakdowns of frontier model architectures and compute markets delivered every weekday. Zero fluff.

🎉 Thank You for Subscribing!

Frequently Asked Questions
Mythos 5 at 1531, Fable 5 at 1525, Opus 5 at 1522. Sol holds 1514, Opus 4.8 1512. Kimi K3 OSS sits exactly at 1500 with Qwen3.8 Max at 1491 and GLM-5.2 at 1483.
31 Elo separate Opus 5 from Qwen3.8 Max at 4x output cost. GLM-5.2 trails Fable 42 points at one-eleventh output price with triple speed. Bulk work cannot justify flagship rates at these gaps.
Interactive to 130+ tok/s models, release-blocking to 1510+ band, bulk to open value tier, rest to standard. Review band floors monthly and ship on your own twenty-ticket deltas.
Lists show Sol at $5/$30 while Reuters reported $4/$20 promo from Aug 21. Budget on list, enjoy promo. Orchestration routing at $2/$6 undercuts both regardless.
Deepak Bagada
Author Profile

Deepak Bagada

Founder & Editor-in-Chief

Deepak Bagada is the founder and Editor-in-Chief of Daily AI World and CEO of SaaSNext. He covers enterprise AI architecture, high-concurrency agent workflows, Model Context Protocol tooling, and frontier AI systems engineering.

Related Intelligence Analysis

Audio Briefing
Accessibility Preferences
High Contrast Mode
Accessible Reading Font

Keyboard Shortcuts

Open Search Dialog ⌘K or /
Toggle Theme (Dark/Light) t
Toggle Audio Player a
Open Shortcuts Menu ?
Close Active Dialog Esc

Cookie & Privacy Preferences

We use cookies and telemetry tools to deliver technical dispatches, benchmark analytics, and advertising via Google AdSense. Review our Privacy Policy.