The Aggregate Digest — 2026-07-22
The Aggregate Digest — 2026-07-22
Daily
New Benchmarks (14)
- LLM Arena RU (Arena Elo): Gemini 3 Pro (Preview) leads with 1210.0 across 92 models. Crowdsourced blind pairwise battles in Russian: Bradley-Terry Elo rating from user votes on llmarena.ru, the Russian-language counterpart of LMArena.
- AA-Briefcase (Rubric Pass Rate (%)): Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) leads with 55.96 across 47 models.
- FutureSim (Top-1 Accuracy (%)): Claude Fable 5 leads with 26.52 across 9 models.
- ProfBench (Overall Rubric Score (%)): Claude Opus 4.7 (Thinking) leads with 61.3 across 98 models. Professional report-writing tasks graded against rubrics written by PhD physicists and chemists and MBA-level finance and consulting experts; rubric satisfaction across the four domains.
- SalesBench (50-Lead Reward): GLM-5.1 leads with 0.508 across 7 models.
- DrugDiscoveryBench (Score (%)): GPT-5.5 leads with 51.6 across 7 models.
- TextQuests (No Clues) (Avg Game Progress (%)): Claude Fable 5 leads with 54.48 across 35 models. Playing 25 classic Infocom interactive-fiction games (Zork etc.) without hints: average game completion via long-horizon exploration, puzzle solving and state tracking in a text world.
- TextQuests (With Clues) (Avg Game Progress (%)): GPT-5.6 Sol leads with 86.84 across 35 models. Playing 25 classic Infocom interactive-fiction games with the official hint books available: average game completion, testing how well models exploit in-context guidance over long horizons.
- RevengeBench (Arena Mean Score): GPT-5 leads with 71.92 across 12 models.
- GameCraft-Bench (Overall Score): Claude Fable 5 leads with 65.7 across 10 models.
- Kagi LLM Benchmark (Accuracy (%)): Claude Fable 5 (Thinking) leads with 91.4 across 145 models. Rotating set of unpublished reasoning, coding and instruction-following tasks run by Kagi; accuracy on ~116 private tasks designed to resist benchmark contamination.
- SpeechMap Compliance (% Requests Completed): sherlock-dash-alpha leads with 100.0 across 344 models. Willingness to answer sensitive or controversial prompts (political argument, satire, religion, rights advocacy): share of requests completed directly rather than refused, evaded or errored. Behavioral propensity, not capability.
- SnitchBench (Govt Contact Rate (%)): Grok 4.1 leads with 92.5 across 8 models. Propensity to autonomously report user wrongdoing: rate at which a model, given evidence of corporate misconduct and email tools, contacts government authorities. Behavioral propensity, not capability.
- MLS-Bench Lite (Score): Claude Fable 5 leads with 49.9 across 12 models. Official 30-task subset of MLS-Bench for evaluating whether AI systems can invent generalizable and scalable machine-learning methods.
New Models (2)
- Gemini 3.6 Flash — ELO 2417, #23/1488, above Muse Spark, below Gemini 3.5 Flash
- RuneBench: 7454.0 (#1/39)
- Design Arena (3D): 1384.0 (#2/133)
- LLM Stats (OSWorld-Verified): 83.0 (#3/21)
- AI Chess Leaderboard (Continuation): 1782.0 (#3/244)
- LLM Stats (CharXiv-R): 89.4 (#5/47)
- Arabic Broad Leaderboard: 9.111 (#6/106)
- AI Chess Leaderboard (Reasoning): 1763.0 (#6/300)
- Kaggle FACTS Grounding: 72.51 (#6/40)
- LLM Stats (MRCR v2 (8-needle)): 54.0 (#7/21)
- Vals AI MedCode: 53.15 (#7/74)
- Gemini 3.5 Flash Lite — ELO 1973, #146/1488, above Qwen 3.5 27B, below GPT-5.1 Codex Mini
- AI Chess Leaderboard (Continuation): 1512.0 (#9/244)
- LLM Stats (OSWorld-Verified): 74.0 (#11/21)
- LLM Stats (Terminal-Bench 2.1): 54.0 (#15/15)
- LLM Stats (MRCR v2 (8-needle)): 21.3 (#18/21)
- RuneBench: 2254.0 (#18/39)
- AA MMMU-Pro: 79.02 (#24/228)
- AI Chess Leaderboard (Reasoning): 1399.0 (#24/300)
- LLM Stats (CharXiv-R): 76.5 (#29/47)
- LLM Stats (GDPval-AA): 1140.0 (#30/41)
- Arabic Broad Leaderboard: 8.626 (#34/106)
New Scores From Top-10 Models (6)
- Claude Opus 4.8 on Guesswork: 0.977 MAE (z-score units) (#12/21)
- GPT-5.6 Pro Sol on MineBench: 2159.91 Elo Rating (#1/51)
- GPT-5.6 Sol on Epoch AI - Vending Bench 2: 9619.37 Score (#2/54)
- GPT-5.6 Sol on HiL-Bench: 32.33 Combined Pass@3 (%) (#8/15)
- GPT-5.6 Sol on Vals AI SkillsBench: 54.1 Accuracy (%) (#9/19)
- Kimi K3 on GDP.pdf: 19.0 Strict Pass Rate (%) (#11/22)
New #1 Leaders (9)
- RuneBench (Total Peak XP Rate (XP/min)): Gemini 3.6 Flash (7454.0) beat GPT-5.6 Sol (7400.0) by 54.0.
- HiL-Bench (Combined Pass@3 (%)): Claude Fable 5 (56.33) beat GPT-5.5 (29.1) by 27.23.
- EnterpriseRAG Bench - Completeness (Score (%)): Troml (81.84) beat OpenClaw (72.86) by 8.98.
- EnterpriseRAG Bench (Overall Score): Troml (76.79) beat OpenClaw (68.22) by 8.57.
- EnterpriseRAG Bench - Recall (Score (%)): Troml (86.55) beat OpenClaw (79.02) by 7.53.
- Vals AI SkillsBench (Accuracy (%)): Grok 4.5 (66.03) beat GPT-5.5 Codex (62.55) by 3.48.
- EnterpriseRAG Bench - Correctness (Score (%)): Troml (83.8) beat OpenClaw (81.6) by 2.2.
- BoxPwnr CTF Bench (Average Platform Completion (%)): Grok 4.5 (xHigh) (56.73) beat GLM-5.1 (55.47) by 1.26.
- LLM Stats (CyberGym) (Score (%)): Gemini 3.5 Flash Cyber (83.2) beat Claude Mythos Preview (83.1) by 0.1.
Don't miss what's next. Subscribe to The Aggregate Digest: