New essay: Stop paying your best model to type
A new piece is up on the site.
I drained a hundred-task prototype under three model policies and priced every task against my own telemetry. Routing a cheaper model — Sonnet on the workers that do most of the typing — while keeping a strong one (Opus) on verification came in at $1.66 a task: about a third of running Opus on everything, at the same 100% first-pass rate. The newest, most expensive model cost the most per task, for a structural reason that has nothing to do with how good it is.
Then the finding turned into a feature: wood-fired-tasks now routes models by pipeline role and by task size.
Don't miss what's next. Subscribe to Stuart Jeff:
Add a comment: