AI broke your delivery dashboard. Measure the review queue instead.

2026-10-05


AI broke your delivery dashboard. Measure the review queue instead.

October 5, 2026 · Issue 115

Deployment frequency and lead time now measure how fast your agents type. The honest constraint moved to human comprehension, and almost nobody is instrumenting it.

For ten years the four DORA metrics were a reasonable proxy for engineering health, because producing a change was the expensive part and everything downstream of it moved at human speed. Faster deploys meant better practices, because better practices were the only way to get faster deploys. That causal chain is what AI quietly cut.

Michael Hill's June piece in LeadDev, "The 8 Software Engineering Metrics AI Broke," makes the case metric by metric. Deployment frequency now signals tool adoption more than engineering maturity. Lead time compresses so far that it stops informing, while the real bottleneck, a human understanding what the machine wrote, hides underneath it. Change failure rate gets murkier still: when an agent writes and a tired reviewer approves, nobody can say whether a failure belongs to the agent, the reviewer, or the process.

The implication for TPMs and tech leaders is uncomfortable. Your existing dashboard probably still looks great. Green charts and a rising throughput line are exactly what you would see in both a healthy AI-assisted org and one quietly accumulating comprehension debt. If the dashboard cannot tell those two worlds apart, it is not a measurement instrument anymore. It is decoration. Today's issue is about what to read instead.


Deep Dive — Metrics & Judgment: Read the Dashboard as a Hypothesis, Not a Verdict

The most useful framing I have found for the AI era comes from the 2026 DORA research, published in May by Google Cloud's DORA team (Nathen Harvey among the authors) with the delta innovation practice, and summarized by InfoQ. Its central claim is that AI is an amplifier: returns come "not from the tools themselves but from a strategic focus on the underlying organizational system." Amplifiers amplify everything. A strong review culture gets stronger; a brittle one fails faster and at higher volume.

Two details in that report matter more than the headline ROI projections. First, the J-curve. The research expects an initial dip in productivity before gains arrive, driven by learning curves, a "verification tax" from reviewing AI-generated code, and downstream processes that have to adapt. If you report only throughput, you will see the early upslope, declare victory, and miss the dip happening somewhere you are not looking. Second, the report links AI adoption with a small but real rise in instability: change failure rates moving from roughly 5% to 6%. One point sounds trivial. At the volume AI enables, it is a lot of extra incidents, and it is the number most likely to be explained away.

So how do you read a dashboard when the classic proxies have decoupled from reality? Treat every metric as a hypothesis about a mechanism, and ask what would have to be true underneath for the number to mean what it appears to mean. Deployment frequency rising is a hypothesis that the team is getting better at safe, small releases. What would falsify it? Review latency growing, review comments shrinking, rollbacks creeping up, on-call load rising. If you are not looking at those, you cannot falsify, and a metric you cannot falsify is a story, not a measurement.

This is where Will Larson's distinction helps. In "Measuring an engineering organization," he argues there are many reasons to measure and each needs a different instrument: measuring to plan, to operate, to optimize, and to inspire. Most dashboards blur all four. Deployment frequency gets used to plan (are we shipping?), operate (are we stable?), optimize (are we improving?), and inspire (look how fast we are!) at once, which is how a single metric ends up carrying four incompatible jobs and failing all of them. AI makes the failure visible because the metric moved while the underlying system did not.

The other half of the answer is balancing a speed measure with something an agent cannot inflate. The DX Core 4, built by Abi Noda and Laura Tacho with input from Nicole Forsgren, Michaela Greiler, and others, pairs Speed (diffs per engineer) with Effectiveness (the survey-based Developer Experience Index), Quality (change failure rate), and Impact (the share of time spent on new capabilities versus maintenance). Notice the design principle: every number is checked by a neighbor. Speed without Quality is volume. Speed without Effectiveness is burnout. Impact is the one that asks whether any of it mattered. Hill's critique is that even these frameworks wobble under AI, since perceived productivity can rise while the ability to assess real output falls. That is fair, and it is exactly why the survey side needs questions about comprehension, not only satisfaction.

What does that look like in practice? Add three instruments that measure the human bottleneck directly. Review wait time and review depth: time a change sits waiting for a human, and a rough signal of how much scrutiny it got. Comprehension spot-checks: periodically ask an engineer to explain, unaided, a recently shipped AI-assisted change. Ownership clarity on incidents: for each postmortem, whether a named human could explain why the code did what it did. None of these is a clean number, and that is the point. Larson's advice to adopt one new metric at a time, and to measure hard things only when they are actionable, applies directly: start with review wait time, which is cheap and tells you whether the bottleneck has moved.

The leadership move is to change what you promise upward. Stop presenting throughput as the headline. Present a pair: the speed number and the number most likely to contradict it. When they diverge, that divergence is the most valuable thing on the page, and you should say so before someone else finds it.

Try this week. Pull review wait time for your busiest team over the last 90 days and plot it next to deployment frequency. If deploys went up and wait time went up too, you have found your real constraint. Bring that single chart to your next leadership review in place of the standard DORA slide.


Method — Larson's Four Reasons to Measure (Will Larson, "Measuring an engineering organization")

What it is. A way to sort metrics by the job they do for you: measure to plan, to operate, to optimize, or to inspire. Each purpose gets its own small set of measures, rather than one dashboard doing everything.

When to use it. When someone asks "what should we measure?" and you suspect the real question is unclear, when a single metric is being used in multiple conversations, or when a metric you trusted has suddenly stopped meaning anything (as AI has done to throughput).

How to run it:

  1. List every metric currently on your dashboard or in your last three leadership updates.
  2. For each, write down which of the four jobs it is actually doing: plan (shipped projects and their impact), operate (incidents, downtime, latency, cost per request), optimize (SPACE-style or survey measures of the engineering system), or inspire (generational improvements that show engineering's business impact).
  3. Flag any metric doing more than one job, and any job with no metric at all. Operate is usually overserved; optimize and inspire usually are not.
  4. Add or retire one metric at a time. Larson's guidance is to measure hard things only if you can act on them, and to use easy measurements to build trust first.
  5. Re-check each quarter whether the metric still means what you assumed. Mechanisms change, and AI just proved it.

When NOT to use it. When you need a fast external answer for a board or CEO question: Larson's own advice there is to reuse your planning or operating metrics rather than inventing new ones.

Example: a team's deployment frequency doubles after adopting coding agents. Under this lens, it is an operate metric being quoted as an optimize success, and the gap shows exactly which measure is missing.


Field Notes

New DORA Report Claims Strong Engineering Foundations Drive AI Return on Investment — The May 2026 DORA research frames AI as an amplifier, models a J-curve with a "verification tax," and ties ROI to platform, version control and data foundations; useful ammunition when finance asks why the gains are slower than the demo.

The 8 Software Engineering Metrics AI Broke — Michael Hill walks through DORA, SPACE and LOC metrics one at a time and explains what each now fails to tell you; a good checklist before your next metrics review.

How DX Core 4 aims to unify developer productivity frameworks — LeadDev's take on combining DORA, SPACE and DevEx into four balanced dimensions; worth reading for the design principle of pairing every speed measure with a counterweight.


Events


Reading


"When a measure becomes a target, it ceases to be a good measure."

— Marilyn Strathern, restating Goodhart's Law (1997)


Don't miss what's next. Subscribe to Critical Path: