The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
August 9, 2026

The open-weights crown holder spent the week taking code arenas off Claude Opus 4.7

The Aggregate Digest — Sunday, August 9, 2026

The open-weights crown holder spent the week taking code arenas off Claude Opus 4.7. Kimi K3 now leads both LMArena WebDev Arena (1682.01) and Arena AI Code (1682.0), and in each case the model it displaced was Claude Opus 4.7, by 114 and 112 points respectively. Those are the two largest leader changes of the week by a wide margin. K3 has held the best-open-weights track since July 27 — thirteen days — and sits sixth overall at 1755 Elo. An open model owning the two head-to-head coding arenas outright is the kind of result that used to take a quarter to arrive.

Claude Fable 5 is the model directly behind it on both boards, at 1630 and 1629.56, and it had a genuinely strong week elsewhere: first on EQ-Bench at 2049.7 across 79 models, first on the SpacetimeDB C# eval at 98.9, and a top-two placing on a dozen more. The honest counterpoint is that its losses came from inside the house. Claude Opus 5 took Design Arena (Game Dev) from it, 1426 to 1394, and Toolathlon by a wider 80.6 to 77.9. Fable 5 spent the week being beaten by a stablemate more often than by anyone else.

Two new benchmark families landed, and both are more interesting for their disagreements than their leaders. Coarena runs live computer-use tasks and lets blind human judges pick the winner. Its Elo board and its task-completion board do not agree: Claude Fable 5 wins the judged rating at 1079.1, while GPT-5.6 Terra finishes the most tasks at 81.4 and does not make the judged top three at all. Doing the most and being judged best are measurably different things here, which is the whole argument for running both tracks.

Featherbench, 28 practical tasks scored by automated checkers, arrived close to saturated: four of its thirteen models — Gemini 3.6 Flash, GPT-5.5, Grok 4.5 and Haiku 4.5 — are tied at 96.0 pass rate. Its companion rubric track, where a judge scores answer quality 0–10, still separates them, and there Kimi K3 leads at 9.5 ahead of Opus 5 at 9.4. A pass rate that four models share is not a ranking; treat the rubric column as the real board.

Elsewhere, the frontier itself did not move. GPT-5.5 Pro holds the best-usable track at 1778, now 108 days; Claude Mythos 5 holds best-measured at 1806, 61 days.

Read the full digest on The Aggregate →

Don't miss what's next. Subscribe to The Aggregate Digest:
← Newer Grok 4.6 came into our main table today at #12, above its rating on agentic work and below it on maths Older → Supabase shipped its agent evals as a matched pair, with and without a skills pack — and the pack helps some models while hurting others
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.