2026-10-06
October 6, 2026 · Issue 116
Engineering sees the token spike in real time but can't act on it; finance can act but sees it weeks late. The fix is organizational, not technical: one accountable name per AI workload.
Every cost model your org has ever used assumes spend scales with headcount. Seats, licenses, cloud reservations: you could forecast them from a hiring plan and a unit price. Agents break that assumption. A token meter runs whether or not the output was useful, and an agent running unattended at 2 a.m. spends exactly like one a human is watching.
IDC's Ryan Smith and Rick Villars put the structural problem sharply in their August piece, "Tokenmaxxing Is Dead. The Governance Gap Isn't." In their words, engineering teams typically spot cost anomalies in real time but lack the budget authority to act on them, while finance holds that authority but often can't detect a problem until a vendor invoice arrives weeks later. Their data points the same direction: 61% of organizations exceeded their 2025 cloud AI budget, blended token prices fell about 50% year over year without shrinking bills (usage outran price), and 83% of IT leaders called token-based pricing the single largest barrier when evaluating AI vendors.
Notice what is not in that list: a tooling gap. The gap is a decision-rights gap. The person who sees the problem cannot decide, and the person who can decide cannot see. That is a classic TPM-shaped hole, and it is why this belongs on your desk rather than only on the CFO's. Below is how I would close it.
The reflex response to runaway AI spend is a cap. Per-engineer ceilings, per-tool limits, a gateway that cuts off at some threshold. Caps are a legitimate floor, and some large companies have reportedly adopted them (a September Gorilla Logic piece cites Uber's per-employee monthly cap and a budget that ran out in four months; treat those figures as secondhand, but the pattern is familiar to anyone who has watched a pilot go viral inside a company). But a cap is a blunt instrument. It cannot tell the agent that is burning tokens on a flaky retry loop from the one that is quietly replacing three days of migration toil. Both hit the ceiling at the same time.
The better frame is a portfolio with owners. Stop thinking of "the AI bill" as one number and decompose it into workloads: the code-review agent, the incident-summarizer, the internal support bot, the migration swarm. Each gets a named, accountable owner before it gets budget. That is IDC's central recommendation, and it is right. Ownership does three things a cap cannot.
First, it reunites seeing and deciding. The owner gets the real-time cost signal and the authority to turn something down. The authority gap IDC describes closes not by giving engineers the credit card but by giving a named person both the dashboard and the pen.
Second, it forces a unit of value before a unit of cost. An owner has to answer "per what?" Per merged change, per resolved ticket, per migrated service. Without a denominator, token counts are vanity: this is exactly the "tokenmaxxing" failure, rewarding burn instead of outcome. Pair this with last week's argument about reading dashboards as hypotheses: cost-per-outcome is a hypothesis about value, and the owner's job is to keep testing it.
Third, it makes variance a conversation instead of a surprise. IDC suggests reviewing token-volume variance quarterly; at agent adoption speeds I would run it monthly for new workloads and quarterly once a workload stabilizes. The point is not the cadence. It is that variance has a standing forum with a named person in the room.
There is a second reason to do this now, and it comes from the velocity side. Gergely Orosz's September reporting on OpenAI's internal "software factory" describes pull request volume rising roughly 10x in six months, with bottlenecks cascading through version control, CI, and deployment. Andrew Ambrosino, desktop lead for Codex, is quoted there: "Everything is now a coding agent. Whether visible code is your output or not, agents write your artifacts." If that trajectory is even directionally right for your org, the cost surface is not just engineers' coding assistants. It is every function that adopts an agent organically, and the same piece notes finance, legal, and recruiting teams adopting without top-down mandates. Your ownership model has to reach beyond engineering or it will have a hole exactly where the next surprise lands.
What does a TPM contribute? Three things nobody else is positioned to do. You can build the inventory, because you already know which teams run what. You can broker the owner assignment, which is a decision-rights negotiation, and you have done a hundred of those. And you can instrument the review, turning the cost-per-outcome number into a recurring artifact rather than a one-time audit. Will Larson makes a related point from the build side in his "Building internal agents" series at Imprint: every company should be doing this work internally, and a prototype by one or two engineers teaches you the requirements faster than a vendor evaluation. Whoever builds the agent should be the first candidate for owning its cost.
One caution on the opposite failure: do not let governance become a gate that kills the experimentation producing the value. Give new workloads a small, explicit sandbox budget with a named owner and an expiry date. Ownership is the requirement; approval is not.
Try this week. List every AI workload your org runs (include the ones in non-engineering teams) in a single table with three columns: owner, cost-per-what, last month's variance from forecast. Any row with an empty owner cell is your first escalation. Bring the table, not an opinion, to your next leadership sync.
What it is. A decision-rights framework that names a Driver (runs the process), an Approver (single person who decides), Contributors (supply input), and Informed (told after). Its core value is that there is exactly one Approver.
When to use it. When a decision has cross-functional stakeholders and an unclear owner. A new AI workload's budget and kill criteria is a textbook case: engineering, finance, and a business function all have a stake, and nobody currently holds the pen.
How to run it:
When NOT to use it. For reversible, low-cost decisions where the overhead of naming roles exceeds the cost of being wrong; just decide and move on.
Example: "DACI for the migration-swarm budget: Driver, platform TPM; Approver, VP Platform; Contributors, FinOps and security; Informed, product directors."
Tokenmaxxing Is Dead. The Governance Gap Isn't — IDC's Smith and Villars name the real failure (budget authority and cost visibility live in different departments) and prescribe one accountable owner per AI workload.
Inside OpenAI's agentic software factory — Pragmatic Engineer's look at what ~10x PR volume does to review, CI, and deployment; the capacity-planning lesson applies well beyond OpenAI.
Why Token Costs Are Breaking the CFO Budgets in 2026 — A vendor-adjacent but useful overview of why headcount-based budgeting fails for agents; read the figures as illustrative, not audited.
"Everything is now a coding agent. Whether visible code is your output or not, agents write your artifacts."
— Andrew Ambrosino, Desktop Lead, OpenAI, as quoted in The Pragmatic Engineer (Sept 15, 2026)
Don't miss what's next. Subscribe to Critical Path: