A Flash model just landed at #3
The Aggregate Digest — Friday, August 14, 2026
A Flash model just landed at #3. Gemini 3.7 Flash enters our main table today at 1780 ELO across 65 boards, one point behind GPT-5.5 Pro and level with Claude Opus 5. The whole top four is separated by five points. Google's own previous best Flash, 3.6, sits 36 points back.
It arrives winning four boards outright: AA MMMU-Pro at 85.49, first of 240; MRCR v2 8-needle at 97.0, taking the lead from GPT-5.6 Sol's 91.5; LVBench at 85.4, from Qwen 3.8 Max's 81.8; and AutomationBench at 30.4, from Claude Opus 5 running at max effort. On AA GPQA Diamond it places second of 575 at 94.55.
Every score also comes with an expectation — the figure a model of that rating usually posts on that board — and this model's gaps from it all lean one way. The board it beats its own rating on by the widest margin is MRCR v2 8-needle, which is long-context retrieval. The boards it misses by the widest margins are all Agent Arena: praise-versus-complaint, bash recovery, steerability. Agents' Last Exam sits below the line too. Long context and multimodal are what carried it to #3; sustained agentic work is not.
The caveat is the sample. 65 boards is thin beside the 220 behind Claude Opus 5 and the 271 behind Claude Fable 5, both within four points of it. The rating is real, but its error bars are wider than those of the models it is sitting next to, and it will move as the boards catch up.
No crown changed today. GPT-5.5 Pro has been the best available model since 23 April.