Your agents don't need a manager. They need a delegation board.

2026-09-29


Your agents don't need a manager. They need a delegation board.

September 29, 2026 · Issue 109

The scarce resource in an AI-heavy org is no longer execution. It is the decision, and most teams have never written down who is allowed to make one.

Read this month's engineering-leadership writing side by side and one pattern shows up. Will Larson is running a "software factory" loop at Imprint. The Pragmatic Engineer reports that OpenAI's agents now carry a change from context-gathering through deploy and incident triage. Duckbill has doubled its weekly pull-request throughput by reviewing only the risky changes. Each of these is described as a tooling story. Each is really a story about delegation.

When an agent can produce a pull request in minutes, the queue that grows is the queue of things waiting for a human to say yes. Review, approval, the architecture call, the "is this safe to ship" call. The teams that pull ahead are not the ones with the best model. They are the ones that decided in advance which yeses can be automated, which need a human, and which need a specific human.

That is a program-management problem before it is an engineering one. It is also the reason TPMs have an unusual opening right now. The org chart says who reports to whom. Nobody owns the map of who may decide what, and agents make the absence of that map expensive very quickly.

Below: what the evidence says about where the bottleneck moved, and a 45-minute workshop format for drawing the map.


Deep Dive — AI as leverage means moving the decision boundary on purpose

The bottleneck moved. The Pragmatic Engineer's look at code review cites GitHub data showing roughly a fivefold rise in pull requests over three years, accelerating from late 2025. One example: a five-person team facing 60 open PRs, about two days of pure review work. When production of change gets cheap, the review queue becomes your throughput ceiling. Adding reviewers does not fix a ceiling that is set by the policy about what needs a reviewer.

The leaders getting leverage are changing the policy, not adding people. Per the same report, Anthropic and OpenAI require human review only for high-risk changes such as authentication, database schemas and APIs. Duckbill moved to that model, tightened its automated guardrails, and went from 80 to 154 PRs merged a week. Notice what was actually done: a risk classification, and a rule attached to each class. That is a delegation decision written down.

OpenAI's factory, as described by the Pragmatic Engineer, makes the same move at every stage: specialized review agents by domain, low-risk changes auto-approved, high-risk changes escalated, deployment agents watching rollouts. The humans' job in that description is judgment, prioritization and oversight of risk.

The same shift is happening in planning. Larson's "Roadmap decisions rather than dates" argues that modern organizations are constrained by decision speed, not execution time. His practical suggestion: pick a plausible external date, then ignore it internally, and instead get teams empowered to make architecture calls locally, with AI-assisted context documents and function-specific harnesses handling routine cross-functional sign-offs. His passkey example shipped without ever appearing on the roadmap, because prototyping behind a feature flag turned abstract tradeoffs into concrete ones fast enough to decide.

His software-factory write-up makes a quieter point that matters for anyone who runs programs. A loop that audits project goals, reads metrics, picks unblocked work and re-runs evaluations only functions if the goals are written explicitly. It removes the option of keeping "how are we doing?" as an unspoken judgment in someone's head. Automating a loop forces you to state the decision rule.

What goes wrong without the map. Two failure modes, both visible in the reporting. First, under-delegation: every agent output routes to the same three senior reviewers, the queue explodes, and leaders conclude that AI "doesn't move the needle." Second, over-delegation by accident: nobody decided that database migrations need a human, and the agent tooling made it easy enough that one shipped without one. The Pragmatic Engineer piece also flags noise: unfiltered AI review output can overwhelm the humans it was meant to help, which is why some companies build filtering layers on top.

Why a TPM should own the exercise. Delegation levels are cross-team by nature. Security has opinions about auth changes. Data has opinions about schemas. Finance has opinions about spend. The TPM is often the only person with standing in all those rooms and no stake in winning any of them. The deliverable is a one-page table: decision types down the side, a delegation level for humans and a separate one for agents across the top, an owner and an escalation path for each row.

Two cautions. The map is a hypothesis, not a policy engraved in stone. Review it after your first serious incident and quarterly otherwise. And do not treat throughput as the scoreboard. A doubling of merged PRs is a leading indicator only if change failure rate and time to restore stay flat. If you would not accept a dashboard that showed only speed, do not accept one here.

Try this week. Pick the one decision type that most often queues behind a human in your program (production config changes, schema migrations, external API changes, whatever it is). Ask its current approver two questions: "What fraction of these do you change or reject?" and "What would have to be true for you to be comfortable not seeing the low-risk ones?" If the rejection rate is near zero, you have a candidate for delegation. Write the answer as a rule with a rollback trigger, and propose it at your next staff meeting.


Method — Delegation Poker (Jurgen Appelo, Management 3.0)

What it is. A card game from Management 3.0 that makes a team's delegation expectations visible by having everyone privately rate, then simultaneously reveal, how much authority a given decision should carry. It uses seven levels: Tell, Sell, Consult, Agree, Advise, Inquire, Delegate.

When to use it. When authority is ambiguous: a new agentic workflow is rolling out, a reorg changed who owns what, or the same decision keeps escalating for no good reason. Also useful when leadership and the team describe the same decision as "delegated" and "not delegated" respectively.

How to run it:

  1. List 6 to 10 concrete decision situations ("merge a change to the payments service," "roll back a release," "approve a model upgrade for the coding agent," "spend above the monthly token budget"). Specific beats abstract.
  2. For each situation, every participant privately picks a card from 1 to 7 for the level they believe is right.
  3. Reveal simultaneously. Per Management 3.0's rules, players score by card value except the "highest minority," a lone high vote that does not count, which discourages grandstanding.
  4. Have the outliers explain their reasoning, then agree on a level for the situation and record it on a Delegation Board: a visible grid of decision areas and their agreed levels.
  5. Run it again for the agent lane: which of those same situations can an agent do at level 7, which at "Consult," and which must stay human. Attach a named owner and an escalation trigger to each.

When NOT to use it. Skip it when the decision is not actually yours to delegate (regulatory sign-offs, legal holds) or when the group has a live trust problem that a game will paper over instead of surface.

Example row: "Approve a low-risk dependency bump: team = Delegate; agent = Delegate with auto-rollback on failed canary. Approve a schema migration: team = Agree; agent = Advise only."


Field Notes

Roadmap decisions rather than dates — Will Larson argues that decision speed, not execution speed, is the constraint, and shows a feature that shipped without ever being a roadmap item. Required reading for anyone whose job is to keep dates honest.

What is happening with code reviews? — The Pragmatic Engineer on the fivefold PR surge and the risk-based review models at Anthropic, OpenAI and Duckbill. The best current evidence that review policy is the lever.

Inside OpenAI's agentic software factory — A stage-by-stage look at how a frontier lab routes work between agents and humans, including what gets auto-approved and what gets escalated. Steal the structure, not the tooling.


Events


Reading


"iterative exploration discarded many of the initial options until the inherent constraints ... simplified the problem."

— Will Larson, Roadmap decisions rather than dates


Don't miss what's next. Subscribe to Critical Path: