The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
August 16, 2026

The model that won the most benchmarks this week is not the one holding the crown

The Aggregate Digest — Sunday, August 16, 2026

The model that won the most benchmarks this week is not the one holding the crown. Claude Opus 5 took the lead on 15 boards over seven days, and no other model took more than four. Gemini 3.7 Flash holds best available model, which it won yesterday.

Both rate 1782. Claude Opus 5 is ahead by less than a point, enough to put it first on our main table and nowhere near enough to move a crown, which needs a five-point gap. Our ranking and our crown name different models this morning, and neither is wrong.

Claude Opus 5 won across subject areas. CHI-Bench went from Claude Opus 4.8's 37.3 to 54.7. Android Bench went to 91.8, MedCode to 63.57, MedScribe to 90.98, ProofBench to 78.0. Arena AI Code at 1692 and LMArena WebDev Arena at 1691.77 both came off Kimi K3, and OpenRouter BrowseComp with search enabled went to 89.0 percent.

Four boards went the other way. Gemini 3.7 Flash took two of them, AutomationBench and AA MMMU-Pro, both running at high reasoning. Kimi K3 took Design Arena (Game Dev) at 1432, and Claude Fable 5 took FrontierCode at 63.6.

The case against Gemini 3.7 Flash is that we have measured it for three days. Its rating rests on 97 benchmarks where Claude Opus 5's rests on 231, so it has further left to move. Of the 90 scores it added this week, second of 169 on Design Arena (Website) at 1348 and fourth of 77 on SimpleQA Verified at 71.2 sit next to 27th of 46 on Multi-turn Debate, a board Claude Opus 5 leads at 1748.5 against Gemini 3.7 Flash's 1476.3.

Neither other crown moved. Claude Mythos 5 has held best measured model for 68 days at 1809. Nobody can buy it, so it never competes for best available model. Kimi K3 has held best open-weights model for 20 days at 1756, and it leads Design Arena (Website) at 1370.

Elsewhere. Grok 4.6 took AA GPQA Diamond at 94.95 from a field of 575. Qwen 3.8 Max scored a flat 100.0 on Coarena Task Completion against a runner-up at 85.9, and a board someone has finished has stopped measuring. ProgramBench sits at the other end, where the best score in the field is Claude Opus 5 resolving 4.5 percent.

Read the full digest on The Aggregate →

Don't miss what's next. Subscribe to The Aggregate Digest:
← Newer GLM-5.3 now has enough scores to rank, and it enters our main table eleventh of 681 Older → Gemini 3.7 Flash takes best available model, ending a 114-day reign
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.