The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
July 21, 2026

AI Benchmark Digest — 2026-07-21

AI Benchmark Digest — 2026-07-21

View on AI Benchmark Hub

Daily

New Benchmarks (10)

  • MathArena - ARXIV_FALSE June (Accuracy (%)): GPT-5.5 (xHigh) leads with 69.44 across 12 models.
  • MathArena - ARXIV June (Accuracy (%)): GPT-5.6 Sol (Max) leads with 86.73 across 12 models.
  • Office Comprehension Bench - File Fidelity Excel (File Fidelity Accuracy (%)): GPT-5.5 (Thinking) leads with 72.6 across 4 models. Perceiving structural and visual content in native .xlsx files — tables, charts, embedded images and formatting — answered from the file itself rather than a text conversion.
  • Office Comprehension Bench - File Fidelity PowerPoint (File Fidelity Accuracy (%)): Claude Opus 4.7 leads with 86.6 across 4 models. Perceiving structural and visual content in native .pptx decks — slide layouts, charts, images and formatting — answered from the file itself.
  • Office Comprehension Bench - File Fidelity Word (File Fidelity Accuracy (%)): Claude Opus 4.7 leads with 91.5 across 4 models. Perceiving structural and visual content in native .docx files — tables, figures and formatting — answered from the file itself; the widest model spread of the three applications.
  • GRIPS (Accuracy (%)): GPT-5.5 (xHigh) leads with 96.4 across 27 models. German Reasoning, Idioms, Puzzles & Wordplay (Sprachspiel): a held-out private set of German-language puzzles, idiomatic expressions and wordplay that resist translation, scored by LLM judge.
  • GRIPS-hard (Accuracy (%)): GPT-5.5 (xHigh) leads with 87.3 across 27 models. The hard subset of GRIPS, restricted to the German reasoning, idiom and wordplay items that discriminate most between models; scores drop sharply versus the full set.
  • Guesswork (MAE (z-score units)): Gemma 3 4B (IT) leads with 6.7079 across 17 models. This site's own live benchmark for predicting benchmark scores: frontier and budget LLMs forecast newly published (model, benchmark) results a day before they land, scored in per-benchmark z-units (MAE, lower is better) against the same-information statistical baseline The Aggregate.
  • StructureClaw (Automatic Workflow) (Success Rate (%)): Kimi K2.6 leads with 100.0 across 9 models. StructureClaw executable agent benchmark (arXiv 2607.14896); the automatic-workflow track reports the success rate of building a working structured-extraction agent without human guidance.
  • LLM GPU Kernel Generation (Pass Rate (%)): GPT-5.5 leads with 100.0 across 5 models. Trace-driven benchmark for LLM-generated GPU kernels (arXiv 2607.14541): the fraction of generated kernels that both compile and produce numerically correct results against reference execution traces.

New Scores From Top-10 Models (36)

  • Claude Opus 4.8 on Mercor APEX: 42.5 Best Pass@1 (%) (#2/49)
  • GPT-5.4 Pro on BenchLM: 60.9 Overall Score (#44/200)
  • GPT-5.5 Pro on BenchLM: 63.7 Overall Score (#38/200)
  • GPT-5.6 Sol on BenchLM: 82.0 Overall Score (#3/200)
  • GPT-5.6 Sol on Icelandic LLM - ARC-Challenge-IS: 95.22 Score (%) (#2/92)
  • GPT-5.6 Sol on Icelandic LLM - Belebele-IS: 94.67 Score (%) (#5/92)
  • GPT-5.6 Sol on Icelandic LLM - GED: 75.5 Score (%) (#10/92)
  • GPT-5.6 Sol on Icelandic LLM - Inflection: 97.0 Score (%) (#6/92)
  • GPT-5.6 Sol on Icelandic LLM - WikiQA-IS: 63.21 Score (%) (#5/92)
  • GPT-5.6 Sol on Icelandic LLM - WinoGrande-IS: 94.49 Score (%) (#10/92)
  • GPT-5.6 Sol on Icelandic LLM Leaderboard - Average: 86.68 Average Score (%) (#8/92)
  • GPT-5.6 Sol on Mercor APEX: 39.9 Best Pass@1 (%) (#5/49)
  • Kimi K3 on AI for Education Pedagogy: 87.76 Accuracy (%) (#31/223)
  • Kimi K3 on AI for Education Pedagogy - Maths: 89.68 Accuracy (%) (#21/223)
  • Kimi K3 on AI for Education Pedagogy - Primary: 92.49 Accuracy (%) (#25/223)
  • Kimi K3 on AI for Education Pedagogy - Science: 89.62 Accuracy (%) (#48/223)
  • Kimi K3 on AI for Education Pedagogy - Secondary: 86.64 Accuracy (%) (#40/223)
  • Kimi K3 on AI for Education Pedagogy - Social studies: 84.55 Accuracy (%) (#51/223)
  • Kimi K3 on AI for Education Pedagogy - Technology: 82.08 Accuracy (%) (#71/223)
  • Kimi K3 on AI for Education SEND: 81.65 Accuracy (%) (#32/215)
  • Kimi K3 on Agent Arena: 9.62 Net Improvement (%) (#8/37)
  • Kimi K3 on Agent Arena - Bash Recovery: 6.41 Bash Recovery (%) (#24/37)
  • Kimi K3 on Agent Arena - Confirmed Success: 14.42 Confirmed Success (%) (#3/37)
  • Kimi K3 on Agent Arena - Praise vs Complaint: 20.62 Praise vs Complaint (%) (#3/37)
  • Kimi K3 on Agent Arena - Steerability: 5.58 Steerability (%) (#23/37)
  • Kimi K3 on Agent Arena - Tool Hallucination: 1.08 Tool Hallucination (%) (#24/37)
  • Kimi K3 on BenchLM: 81.0 Overall Score (#4/200)
  • Kimi K3 on Benchmarks.bio - VariantBench: 32.8 Pass Rate (%) (#9/12)
  • Kimi K3 on Creative Writing (Lechmazur): 2.9 Mean Score (#3/39)
  • Kimi K3 on GBA Eval: 0.4835 Overall Emulator Score (#8/23)
  • Kimi K3 on GBA Eval: 0.4835 Overall Emulator Score (#8/23)
  • Kimi K3 on MineBench: 1756.72 Elo Rating (#7/51)
  • Kimi K3 on Multi-turn Debate (Lechmazur): 1740.8 Bradley-Terry Rating (#2/41)
  • Kimi K3 on PrinzBench: 47.0 Score (x/99) (#12/29)
  • Kimi K3 on Riemann-bench: 37.6 Score (%) (#7/21)
  • Kimi K3 on TaxCalcBench: 6.0 Correct Returns - Strict (%) (#9/11)

New #1 Leaders (4)

  • Design Arena (UI Components) (Elo): Kimi K3 (1443.0) beat GPT-5.6 Sol (1388.0) by 55.0.
  • Riemann-bench (Score (%)): GPT-5.6 Sol (Max) (74.4) beat Claude Fable 5 (60.0) by 14.4.
  • PrinzBench (Score (x/99)): GPT-5.6 Pro Sol (91.0) beat GPT-5.5 Pro (Extended) (82.0) by 9.0.
  • OpenClawProBench (Overall Score (%)): Kimi K3 (81.6) beat GLM-5.2 (81.3) by 0.3.
Don't miss what's next. Subscribe to The Aggregate Digest:
← Newer The Aggregate Digest — 2026-07-22 Older → AI Benchmark Digest — 2026-07-20
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.