Harrison Guo logo

Harrison Guo

Archives
Log in
Subscribe
September 14, 2026

Same model, 13.3% to 38.3%

Nothing about the model changed. The score almost tripled.

Same Model, 13.3% to 38.3%

OpenAI changed two API settings on ARC-AGI-3. Same model, same weights, same task set. The score went from 13.3% to 38.3%, and it spent one sixth the output tokens getting there.

The official harness had been throwing away the model's reasoning after every move. So the model kept re-deriving what it already knew, and that rebuilding is what the output tokens were buying. The expensive configuration was expensive because it was worse.

Which makes 13.3% a true fact about a harness that got read as a fact about a model.

Also this week

  • 550 Tokens of Prompt, 588 Tokens of Tools: Pi is famous for a system prompt under a thousand tokens. It is 550. The four tool schemas that ship beside it in every request are 588.

Harrison

Don't miss what's next. Subscribe to Harrison Guo:
Older → The approval prompt is not the sandbox
Powered by Buttondown, the easiest way to start and grow your newsletter.