Grok 4.6 came into our main table today at #12, above its rating on agentic work and below it on maths
The Aggregate Digest — Thursday, August 13, 2026
Grok 4.6 came into our main table today at #12, above its rating on agentic work and below it on maths. The model scored 1745 ELO across 72 benchmarks, which puts it between Qwen 3.8 Max and Gemini 3.1 Pro (Preview). Grok 4.5 sits four places lower and eleven points back, with 215 boards to its name.
Every score also comes with an expectation, the figure a model of that rating usually posts on that board, and Grok 4.6's gaps from it lean one way. Agentic work sits furthest above the line, with planning and code behind it, while maths falls further below than any other skill we track. Its skill scores tell the same story, with agentic at the top and vision-language at the bottom.
The model leads two boards outright. CursorBench 3.1 scores models on ambiguous, multi-file coding jobs taken from real Cursor sessions, and Grok 4.6 tops it at 70.8, while on AA GPQA Diamond it comes first of 570 rows at 94.95.
Both wins come with conditions attached. Each was set at the model's highest effort setting, and both margins are thin, beating Fable 5 Max by 0.3 on CursorBench and GPT-5.6 Sol by 0.8 on GPQA Diamond. At ordinary effort it drops to fourth on CursorBench. Its 72 boards are a small sample next to the 1,047 behind Gemini 3.1 Pro (Preview), which sits one place below it, so the rating is real but its error bars run wider than those around it.
On the 4,880 boards nobody has run it on, our predicted scores put Grok 4.6 at 98.5 on AA TAU-2 Bench and 63.0 on AA Terminal-Bench Hard. Both are agentic tasks, which is where its measured scores are strongest.
Use it at high effort for agent and tool work and for multi-file code, and treat the two board wins as ties. For maths, its own scores point somewhere else.