The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
July 18, 2026

AI Benchmark Digest — 2026-07-18

AI Benchmark Digest — 2026-07-18

View on AI Benchmark Hub

Daily

New Benchmarks (10)

  • Sense Ruby Navigation - Baseline (Mean cited recall across repositories (%)): GPT-5.5 leads with 71.46 across 5 models. Sense Ruby Navigation measures how completely coding agents locate must-find code across 13 pinned Ruby and Rails repositories; this board averages exact path-and-line cited recall for the normal-tools baseline arm.
  • SciDiagramEdit - Semantic Faithfulness (No-skill pairwise win rate (%)): GPT-5.5 leads with 74.5 across 4 models. SciDiagramEdit measures instruction-driven editing of scientific figures reconstructed from real paper revision pairs; this board reports blind pairwise semantic-faithfulness win rate for the unassisted model backbones.
  • SciDiagramEdit - Aesthetic Preference (No-skill pairwise win rate (%)): GPT-5.5 leads with 46.6 across 4 models. SciDiagramEdit measures instruction-driven editing of scientific figures reconstructed from real paper revision pairs; this board reports blind pairwise aesthetic preference for the unassisted model backbones.
  • SVLAT - Overall (Accuracy (%)): Gemini 3.1 Pro (Preview) leads with 88.57 across 6 models. Scientific Visualization Literacy Assessment Test (SVLAT) evaluates reading and reasoning over static and animated scientific visualizations; overall accuracy is averaged across the authors\
  • SVLAT - Image (Accuracy (%)): Gemini 3.1 Pro (Preview) leads with 90.86 across 6 models. Scientific Visualization Literacy Assessment Test (SVLAT) multiple-choice questions about static scientific visualizations; accuracy is averaged across the authors\
  • SVLAT - Animation (Accuracy (%)): Gemini 3.1 Pro (Preview) leads with 82.86 across 6 models. Scientific Visualization Literacy Assessment Test (SVLAT) multiple-choice questions about animated scientific visualizations; accuracy is averaged across the authors\
  • Semgrep IDOR - Multimodal F1 (F1 score (%)): GPT-5.6 Sol leads with 61.7 across 6 models. Semgrep researcher-reviewed benchmark of insecure direct object reference detection in representative production applications over ten runs, using the Semgrep Multimodal security workflow and reporting F1.
  • Semgrep IDOR - Guided Prompt F1 (F1 score (%)): GPT-5.6 Sol leads with 38.9 across 5 models. Semgrep researcher-reviewed benchmark of insecure direct object reference detection in representative production applications over ten runs, using each LLM alone with a guided prompt and reporting F1.
  • Frontier Security Agents - Cybench Hard (Success rate (%)): GPT-5.5 leads with 94.1 across 8 models. Cybench hard-variant offensive-security evaluation of agents solving capture-the-flag challenges, reporting raw success rate at each model\
  • Frontier Security Agents - BOTS v1 (BOTS score (%)): Claude Opus 4.8 leads with 93.9 across 8 models. BOTS v1 defensive-security agent evaluation on scored SOC investigation questions in Splunk, reporting official point-weighted score at each model\

New Models (1)

  • Inkling — ELO 2630, #11/1475, above Claude Opus 4.8, below GPT-5.6 Pro Sol
    • EQ-Bench Longform Writing: 72.5 (#22/123)
    • RuneBench: 1590.0 (#24/33)
    • ARC-AGI-2: 36.53 (#47/166)
    • ARC-AGI-1: 79.5 (#48/164)

New Scores From Top-10 Models (35)

  • Claude Opus 4.8 on BridgeBench Security: 1001.5 Confidence-weighted Elo (#2/5)
  • GPT-5.6 Sol on AI for Education Pedagogy: 91.21 Accuracy (%) (#3/221)
  • GPT-5.6 Sol on AI for Education Pedagogy - Maths: 92.86 Accuracy (%) (#4/221)
  • GPT-5.6 Sol on AI for Education Pedagogy - Primary: 94.37 Accuracy (%) (#7/221)
  • GPT-5.6 Sol on AI for Education Pedagogy - Science: 92.35 Accuracy (%) (#17/221)
  • GPT-5.6 Sol on AI for Education Pedagogy - Secondary: 90.09 Accuracy (%) (#3/221)
  • GPT-5.6 Sol on AI for Education Pedagogy - Social studies: 90.91 Accuracy (%) (#4/221)
  • GPT-5.6 Sol on AI for Education Pedagogy - Technology: 88.68 Accuracy (%) (#5/221)
  • GPT-5.6 Sol on AI for Education SEND: 87.61 Accuracy (%) (#2/213)
  • GPT-5.6 Sol on AI for Education Visual Reasoning: 84.9 Accuracy (%) (#2/64)
  • GPT-5.6 Sol on AI for Education Visual Reasoning - match (figure): 77.0 Accuracy (%) (#3/64)
  • GPT-5.6 Sol on AI for Education Visual Reasoning - pattern completion (2d): 82.2 Accuracy (%) (#3/64)
  • GPT-5.6 Sol on AI for Education Visual Reasoning - reasoning by analogy: 86.5 Accuracy (%) (#2/64)
  • GPT-5.6 Sol on Agent Security League - Functional Correctness: 70.9 Functional Correctness (%) (#11/18)
  • GPT-5.6 Sol on Agent Security League - Security Correctness: 23.5 Security Correctness (%) (#3/18)
  • Inkling on ARC-AGI-1: 79.5 Accuracy (%) (#48/164)
  • Inkling on ARC-AGI-2: 36.53 Accuracy (%) (#47/166)
  • Inkling on EQ-Bench Longform Writing: 72.5 Writing Score (0-100) (#22/123)
  • Inkling on RuneBench: 1590.0 Total Peak XP Rate (XP/min) (#24/33)
  • Inkling on SvelteBench: 87.8 Average pass@1 (%) (#76/151)
  • Kimi K3 on BridgeBench Debugging: 1066.7 Confidence-weighted Elo (#1/3)
  • Kimi K3 on BridgeBench Hallucination: 1030.4 Confidence-weighted Elo (#1/3)
  • Kimi K3 on BridgeBench Refactoring: 1065.7 Confidence-weighted Elo (#1/3)
  • Kimi K3 on BridgeBench Security: 1023.1 Confidence-weighted Elo (#1/5)
  • Kimi K3 on DeepSWE: 68.5 Pass@1 (%) (#5/20)
  • Kimi K3 on EQ-Bench Longform Writing: 79.6 Writing Score (0-100) (#6/123)
  • Kimi K3 on RuneBench: 1891.0 Total Peak XP Rate (XP/min) (#21/33)
  • Kimi K3 on Surface Evolver Bench: 95.0 Mean Score (%) (#2/23)
  • Kimi K3 on Surface Evolver Bench Pass Rate: 87.5 Pass Rate (%) (#2/23)
  • Kimi K3 on Vals AI CyberBench: 79.03 Accuracy (%) (#4/16)
  • Kimi K3 on Vals AI Excel Modeling: 66.4 Accuracy (%) (#5/23)
  • Kimi K3 on Vals AI Finance Agent v2: 54.36 Accuracy (%) (#5/37)
  • Kimi K3 on Vals AI ProofBench: 70.0 Accuracy (%) (#5/50)
  • Kimi K3 on Vals AI SWE-bench Verified: 93.4 Resolved (%) (#3/71)
  • Kimi K3 on Vals AI Vibe Code Bench: 71.06 Accuracy (%) (#8/73)

New #1 Leaders (9)

  • ExploitGym (Successful Intended Exploits (#)): GPT-5.6 Sol (293.0) beat Claude Mythos Preview (157.0) by 136.0.
  • AI for Education Visual Reasoning - match (process) (Accuracy (%)): GPT-5.6 Sol (88.9) beat Gemini 3 Flash (77.8) by 11.1.
  • AI for Education Visual Reasoning - odd one out (Accuracy (%)): GPT-5.6 Sol (88.3) beat Gemini 3.5 Flash (80.5) by 7.8.
  • LLM Stats (MCP Atlas) (Score (%)): Muse Spark 1.1 (88.1) beat Kimi K3 (84.2) by 3.9.
  • LLM Stats (Toolathlon) (Score (%)): Muse Spark 1.1 (75.6) beat Kimi K3 (73.2) by 2.4.
  • AI for Education Visual Reasoning - pattern completion (linear) (Accuracy (%)): GPT-5.6 Sol (92.3) beat Gemini 3.5 Flash (91.5) by 0.8.
  • Terminal-Bench 2.1 (Claude Code) (Accuracy (%)): Claude Fable 5 (83.8) beat Claude 5 Fable (83.1) by 0.7.
  • Terminal-Bench 2.1 (Accuracy (%)): Claude Fable 5 (83.8) beat GPT-5.5 (83.4) by 0.4.
  • CADGenBench (Aggregate CAD Score): GPT-5.6 Sol (xHigh) (0.5319) beat Archie in Forge (0.51) by 0.02.
Don't miss what's next. Subscribe to The Aggregate Digest:
← Newer AI Benchmark Digest — 2026-07-19 Older → AI Benchmark Digest — 2026-07-17
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.