Claude Fable 5.1 arrived today and went straight to the top of the best available model table, ending Claude Opus 5's twelve-day hold at 1809 to 1786
The Aggregate Digest — Wednesday, September 2, 2026
Claude Fable 5.1 arrived today and went straight to the top of the best available model table, ending Claude Opus 5's twelve-day hold at 1809 to 1786. A 23-point margin is wide for this table. The two crown changes before it were worth 6 points and 8.
Claude Fable 5.1 earned that rating on the boards it walked onto. FrontierMath Tiers 1-3 (v2) sets 285 private problems that run from advanced undergraduate mathematics up to early-career research. Claude Fable 5.1 scores 90.18 percent there and is first of 100. GPT-5.6 Sol is second at 89.12. CursorBench 3.1 is Cursor's benchmark of ambiguous multi-file tasks taken from real development sessions. Claude Fable 5.1 takes 73.4 percent and is first of 50. The best entry that is not a Claude is Grok 4.6 at 70.8. AA-Briefcase asks a model to do knowledge work against a set of real files. Claude Fable 5.1 leads its 31 entries at 61.52 percent. The best Claude Opus 5 entry there is 57.98. LLM Stats runs OSWorld 2.0, a suite of long desktop workflows. Claude Fable 5.1 reaches 77.9 and Claude Opus 5 reaches 70.6.
Neither of the other two crowns changed hands. Kimi K3 is still the best open-weights model. Claude Mythos 5 is still the best measured model, now two points behind Claude Fable 5.1 and too close to call a change.
Two places it does not move. FrontierMath Tier 4 (v2) holds out the research-level problems for a separate board. Claude Fable 5.1 scores 87.8 percent there and ties Claude Fable 5. GDP.pdf is new today and scores models on professional PDF pages that keep their working layout. It has no Claude Fable 5.1 row yet. GPT-5.6 Sol leads it at 30.7 percent strict pass and none of its 34 entries clears 31.