Same model, 13.3% to 38.3%
Nothing about the model changed. The score almost tripled.
OpenAI changed two API settings on ARC-AGI-3. Same model, same weights, same task set. The score went from 13.3% to 38.3%, and it spent one sixth the output tokens getting there.
The official harness had been throwing away the model's reasoning after every move. So the model kept re-deriving what it already knew, and that rebuilding is what the output tokens were buying. The expensive configuration was expensive because it was worse.
Which makes 13.3% a true fact about a harness that got read as a fact about a model.
Also this week
- 550 Tokens of Prompt, 588 Tokens of Tools: Pi is famous for a system prompt under a thousand tokens. It is 550. The four tool schemas that ship beside it in every request are 588.
Harrison
Don't miss what's next. Subscribe to Harrison Guo: