The Aggregate Digest logo

The Aggregate Digest

Archives
Log in
Subscribe
August 20, 2026

GLM-5.3 enters the best-available ranking tied for 11th of 681, and its strongest results are on agent benchmarks

The Aggregate Digest — Thursday, August 20, 2026

GLM-5.3 enters the best-available ranking tied for 11th of 681, and its strongest results are on agent benchmarks. A reader who runs agents has a concrete reason to try it: in LLM Stats' testing, GLM-5.3 is first on AutomationBench, Zapier's business-workflow benchmark, ahead of Gemini 3.7 Flash and Claude Opus 5, and first on CyberGym, which has agents find and exploit software vulnerabilities. On FrontierSWE's 20-hour coding-agent tasks it places second of 17, behind Claude Fable 5. One caveat before switching: GLM-5.3 is API-only, while GLM-5.2 before it is open-weights.

The number behind the debut is 1753 ELO over 43 benchmarks. GLM-5.3 arrived provisional on 15 August at 1772 over 13 boards, and the thirty boards measured since then cost it 19 points. The 11th place is shared with Gemini 3.6 Flash, which also rates 1753 with 182 boards behind its number, so GLM-5.3's rating is the one more likely to move from here. Best available model stays Gemini 3.7 Flash at 1787, 34 points above the pair of them.

Z.ai published six boards with the GLM-5.3 launch, and its model comes third on five of them, behind pairs drawn from GPT-5.6 Sol, Claude Fable 5, Claude Opus 4.8 and Kimi K3. Its one win is GDPval-AA v2, which Z.ai scores in a board-local Elo, ahead of Claude Fable 5 and GPT-5.6 Sol. On Terminal Bench 3.0, GLM-5.3 takes 28.3 percent to GPT-5.6 Sol's 34.6. A vendor deck that mostly places its own model third is rare, and the third places match what our fused rating says. GLM-5.2 finishes last on four of the six launch boards and sits 32nd overall.

Read the full digest on The Aggregate →

Don't miss what's next. Subscribe to The Aggregate Digest:
← Newer Claude Opus 5 is now the best available model, after winning two agentic boards in this morning's Artificial Analysis refresh Older → GLM-5.3 now has enough scores to rank, and it enters our main table eleventh of 681
aibenchmarks.dev
Twitter
Telegram
Powered by Buttondown, the easiest way to start and grow your newsletter.