The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
July 30, 2026

Kimi K3 walked onto Epoch AI's hardest boards and placed third on the first one

The Aggregate Digest — Thursday, July 30, 2026

Kimi K3 walked onto Epoch AI's hardest boards and placed third on the first one. On Epoch AI - Scicode it scored 58.68 against a field of 102, behind only Claude Fable 5's 60.19 and Gemini 3.1 Pro's 58.91 — 1.51 points off a crown, on a board where the top four are separated by less than two. Epoch AI - Proofbench put it fifth of 50 at 70.0, with Claude Opus 5 leading on 78.0 and GPT-5.6 Sol and Claude Fable 5 tied on 77.0. Blueprint Bench 2 took it eighth of 21. For a model sitting tenth overall, three top-ten finishes in a single morning is a strong opening, and the honest counterweight is on the same source: Epoch AI - ECI ranks it 59th of 442 at 155.59.

Claude Opus 5 extended a crown streak into a second day. After five yesterday it took three more: Vending-Bench 2, where it banked $11,181.87 against Claude Opus 4.7's $10,936.76; WebDev Arena, at 1711.88 over Claude Fable 5's 1653.93 — and Kimi K3 slid in behind it at 1681.75, so the previous leader is now third; and the new Epoch AI - Mystery Game Puzzles board at 59.0. The fourth overall model has now spent two consecutive days taking territory from everyone above and below it.

Claude Fable 5, seventh overall, answered on Epoch AI - Enigma Eval and answered hard: 39.28, against a standing leader mark of 23.82 from GPT-5.4 Pro. That is a 15.46-point jump in one step, and GPT-5.6 Sol's 37.12 is now the only score within ten points of it. Fable 5 also still holds Blueprint-Bench 2 at 0.386, where Opus 5 and Kimi K3 entered seventh and eighth.

Seven new benchmarks. The pick is Epoch AI - Mystery Game Puzzles, whose whole design is a refusal to say what it tests: models get mid-game positions from a deliberately undisclosed board game, one best move each, and the game's identity is withheld precisely to stop anyone training for it. Twenty models have run it.

Elsewhere. Five Roboflow Playground arena boards landed, and the overall crown there did not go to a language model at all — SAM 3 tops the 62-entrant pooled Elo at 1409, while Gemini 3 Flash takes detection and classification. On Roboflow Vision Evals - Visual Understanding, Gemini 3.5 Flash leads 77 models at 79.1%.

Read the full digest on The Aggregate →

Don't miss what's next. Subscribe to The Aggregate Digest:
← Newer Claude Opus 5 took a fourth straight day of crowns, and its best one was for refusing to answer Older → Claude Opus 5 nearly doubled its own GPU-kernel record overnight
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.