The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
July 22, 2026

The Aggregate Digest — 2026-07-22

The Aggregate Digest — 2026-07-22

View on The Aggregate

Daily

New Benchmarks (14)

  • LLM Arena RU (Arena Elo): Gemini 3 Pro (Preview) leads with 1210.0 across 92 models. Crowdsourced blind pairwise battles in Russian: Bradley-Terry Elo rating from user votes on llmarena.ru, the Russian-language counterpart of LMArena.
  • AA-Briefcase (Rubric Pass Rate (%)): Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) leads with 55.96 across 47 models.
  • FutureSim (Top-1 Accuracy (%)): Claude Fable 5 leads with 26.52 across 9 models.
  • ProfBench (Overall Rubric Score (%)): Claude Opus 4.7 (Thinking) leads with 61.3 across 98 models. Professional report-writing tasks graded against rubrics written by PhD physicists and chemists and MBA-level finance and consulting experts; rubric satisfaction across the four domains.
  • SalesBench (50-Lead Reward): GLM-5.1 leads with 0.508 across 7 models.
  • DrugDiscoveryBench (Score (%)): GPT-5.5 leads with 51.6 across 7 models.
  • TextQuests (No Clues) (Avg Game Progress (%)): Claude Fable 5 leads with 54.48 across 35 models. Playing 25 classic Infocom interactive-fiction games (Zork etc.) without hints: average game completion via long-horizon exploration, puzzle solving and state tracking in a text world.
  • TextQuests (With Clues) (Avg Game Progress (%)): GPT-5.6 Sol leads with 86.84 across 35 models. Playing 25 classic Infocom interactive-fiction games with the official hint books available: average game completion, testing how well models exploit in-context guidance over long horizons.
  • RevengeBench (Arena Mean Score): GPT-5 leads with 71.92 across 12 models.
  • GameCraft-Bench (Overall Score): Claude Fable 5 leads with 65.7 across 10 models.
  • Kagi LLM Benchmark (Accuracy (%)): Claude Fable 5 (Thinking) leads with 91.4 across 145 models. Rotating set of unpublished reasoning, coding and instruction-following tasks run by Kagi; accuracy on ~116 private tasks designed to resist benchmark contamination.
  • SpeechMap Compliance (% Requests Completed): sherlock-dash-alpha leads with 100.0 across 344 models. Willingness to answer sensitive or controversial prompts (political argument, satire, religion, rights advocacy): share of requests completed directly rather than refused, evaded or errored. Behavioral propensity, not capability.
  • SnitchBench (Govt Contact Rate (%)): Grok 4.1 leads with 92.5 across 8 models. Propensity to autonomously report user wrongdoing: rate at which a model, given evidence of corporate misconduct and email tools, contacts government authorities. Behavioral propensity, not capability.
  • MLS-Bench Lite (Score): Claude Fable 5 leads with 49.9 across 12 models. Official 30-task subset of MLS-Bench for evaluating whether AI systems can invent generalizable and scalable machine-learning methods.

New Models (2)

  • Gemini 3.6 Flash — ELO 2417, #23/1488, above Muse Spark, below Gemini 3.5 Flash
    • RuneBench: 7454.0 (#1/39)
    • Design Arena (3D): 1384.0 (#2/133)
    • LLM Stats (OSWorld-Verified): 83.0 (#3/21)
    • AI Chess Leaderboard (Continuation): 1782.0 (#3/244)
    • LLM Stats (CharXiv-R): 89.4 (#5/47)
    • Arabic Broad Leaderboard: 9.111 (#6/106)
    • AI Chess Leaderboard (Reasoning): 1763.0 (#6/300)
    • Kaggle FACTS Grounding: 72.51 (#6/40)
    • LLM Stats (MRCR v2 (8-needle)): 54.0 (#7/21)
    • Vals AI MedCode: 53.15 (#7/74)
  • Gemini 3.5 Flash Lite — ELO 1973, #146/1488, above Qwen 3.5 27B, below GPT-5.1 Codex Mini
    • AI Chess Leaderboard (Continuation): 1512.0 (#9/244)
    • LLM Stats (OSWorld-Verified): 74.0 (#11/21)
    • LLM Stats (Terminal-Bench 2.1): 54.0 (#15/15)
    • LLM Stats (MRCR v2 (8-needle)): 21.3 (#18/21)
    • RuneBench: 2254.0 (#18/39)
    • AA MMMU-Pro: 79.02 (#24/228)
    • AI Chess Leaderboard (Reasoning): 1399.0 (#24/300)
    • LLM Stats (CharXiv-R): 76.5 (#29/47)
    • LLM Stats (GDPval-AA): 1140.0 (#30/41)
    • Arabic Broad Leaderboard: 8.626 (#34/106)

New Scores From Top-10 Models (6)

  • Claude Opus 4.8 on Guesswork: 0.977 MAE (z-score units) (#12/21)
  • GPT-5.6 Pro Sol on MineBench: 2159.91 Elo Rating (#1/51)
  • GPT-5.6 Sol on Epoch AI - Vending Bench 2: 9619.37 Score (#2/54)
  • GPT-5.6 Sol on HiL-Bench: 32.33 Combined Pass@3 (%) (#8/15)
  • GPT-5.6 Sol on Vals AI SkillsBench: 54.1 Accuracy (%) (#9/19)
  • Kimi K3 on GDP.pdf: 19.0 Strict Pass Rate (%) (#11/22)

New #1 Leaders (9)

  • RuneBench (Total Peak XP Rate (XP/min)): Gemini 3.6 Flash (7454.0) beat GPT-5.6 Sol (7400.0) by 54.0.
  • HiL-Bench (Combined Pass@3 (%)): Claude Fable 5 (56.33) beat GPT-5.5 (29.1) by 27.23.
  • EnterpriseRAG Bench - Completeness (Score (%)): Troml (81.84) beat OpenClaw (72.86) by 8.98.
  • EnterpriseRAG Bench (Overall Score): Troml (76.79) beat OpenClaw (68.22) by 8.57.
  • EnterpriseRAG Bench - Recall (Score (%)): Troml (86.55) beat OpenClaw (79.02) by 7.53.
  • Vals AI SkillsBench (Accuracy (%)): Grok 4.5 (66.03) beat GPT-5.5 Codex (62.55) by 3.48.
  • EnterpriseRAG Bench - Correctness (Score (%)): Troml (83.8) beat OpenClaw (81.6) by 2.2.
  • BoxPwnr CTF Bench (Average Platform Completion (%)): Grok 4.5 (xHigh) (56.73) beat GLM-5.1 (55.47) by 1.26.
  • LLM Stats (CyberGym) (Score (%)): Gemini 3.5 Flash Cyber (83.2) beat Claude Mythos Preview (83.1) by 0.1.
Don't miss what's next. Subscribe to The Aggregate Digest:
← Newer Hand a model image tools and its score on the same exam multiplies eight-and-a-half-fold Older → AI Benchmark Digest — 2026-07-21
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.