The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
July 29, 2026

Claude Opus 5 nearly doubled its own GPU-kernel record overnight

The Aggregate Digest — Wednesday, July 29, 2026

Claude Opus 5 nearly doubled its own GPU-kernel record overnight. Yesterday it took KernelBench Hub - CUDA at 103.67% of hardware roofline; today the board reads 196.1, with Claude Fable 5 a distant second on 72.66 and the remaining four entrants all under 43. That metric is a percentage of a modelled hardware ceiling rather than a capped score, so clearing 100 is possible — but clearing 196 means the winning kernel is doing far less work than the reference implementation, which is impressive and worth a second look in equal measure. On the sibling Mega board it leads outright too, 24.29x speedup against Fable 5's 19.12. On Hard, the one KernelBench board where the ceiling still bites, it is second at 52.61 behind Fable 5's 55.21.

The #7 model overall had a broad day beyond kernels: fourteen fresh top-ten placements and four of the day's five new crowns. It took AI for Education Pedagogy - Technology off Kimi K2.5, 91.51 to 89.62, and AI for Education SEND off Gemini 3.6 Flash by less than half a point. Its Max-effort variant claimed both Agent Arena boards that moved, Confirmed Success at 17.56% and Praise vs Complaint at 25.63%.

The same education suite supplies the counterweight. On the parent AI for Education Pedagogy board Opus 5 is only fourth at 91.1, behind GPT-5.5, Gemini 3.1 Pro and GPT-5.6 Sol, and on the Social studies slice it sits 70th of 227 — nine points adrift of the o3 / GPT-5.5 tie at 91.82. Winning the specialist splits while losing the aggregate is a familiar shape.

PerceptionBench arrives from Moonshot AI with ten sub-scores over sixteen multimodal models, and it is emphatically not a one-model board. GPT-5.6 Sol, #10 overall, takes the headline 59.7 and six of the ten slices, topping out at 76.7 on Localization. Seed 2.1 Pro wins OCR (66.7) and Fine-Grained Recognition (56.6), Gemini 3.1 Pro (Preview) takes Contextual Integration, and Gemini 3.5 Flash leads perception hallucination at 50.6. Nothing anywhere on the suite clears 77 — atomic visual perception is nowhere near saturated.

Elsewhere. LLM-SoccerArena opens with GPT-5.5 on 168 prediction points from seven forecasters, Claude Opus 4.8 (156) and DeepSeek V4 Pro (155) trailing. AgentBattler also debuts, but all three entrants are GPT-5.6 harness variants, so read it as an internal ablation rather than a field.

Read the full digest on The Aggregate →

Don't miss what's next. Subscribe to The Aggregate Digest:
← Newer Kimi K3 walked onto Epoch AI's hardest boards and placed third on the first one Older → Claude Opus 5 took five more crowns today, a fourth straight day of collecting — and three of the five came off its own stablemate
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.