The phone-agent race got its own scoreboard today — and the winner isn't who the overall rankings would pick
The Aggregate Digest — Sunday, August 2, 2026
The phone-agent race got its own scoreboard today — and the winner isn't who the overall rankings would pick. MobileWorld, a smartphone-environment benchmark from Tongyi-MAI, lands as split boards. On the GUI-only split — 117 tasks solved end-to-end from screen interaction alone — Kimi K3 takes the crown at 74.4%, clear of GPT-5.6 Sol at 70.1%, even though Sol sits three places above it on the overall board (#8 vs #11). The user-interaction split, 44 tasks that force the agent to stop and ask a simulated user for clarification mid-task, turns the table further: Seed 2.0 Pro — #55 overall — leads at 61.4%, ahead of Claude Opus 4.7 at 59.1%. Tapping through a phone and knowing when to ask a human are, on today's evidence, skills the general leaderboard only loosely predicts.
The day's most sobering debut is InfoOps Bench: 36 frontier models measured on whether they can be co-opted for state-backed information operations, with prompts drawn from a monitoring pipeline tracking Russian, Chinese and Iranian state-backed assets. The Integrity Score is simply the percentage of those requests a model refuses — and the top four slots are all Anthropic: Claude Opus 4.8 at 96.7%, Opus 4.7 at 95.1%, Sonnet 5 at 93.5% and Haiku 4.5 at 91.2%.
ATM-Bench Hard rounds out the debuts: adversarially selected long-term memory questions, answered with the full interaction history provided. It's an OpenAI podium — GPT-5.5 at 71.49, GPT-5 at 66.05, GPT-5.4 at 65.19 — with Gemini 2.5 Pro the closest challenger at 64.3.
Elsewhere. MLE-Bench Lite added its Codex-harness split, a two-model board where GPT-5.5's 68.18% medal average leads MiniMax-M3. And the day's honest counterpoint: Claude Opus 5, #4 overall, managed just 14.6% sign accuracy on HieroglyphBench — 15th of 20 on a board Gemini 3.5 Flash leads at 52.5%. Reading ancient Egyptian, it turns out, is still a specialist's game.