Dispatch

Archives
Log in
Subscribe
April 3, 2026

Your agent pilot is probably going to fail. Here's why it's not the model.

Editor's note

Welcome to Dispatch. One issue a week. No hype. If something in here is wrong, reply and tell me — I'll correct it publicly.


1. The Agent Pilot Graveyard

78% of enterprises have an active AI agent pilot. 14% have reached production scale — defined as handling more than half of the target task volume with automated quality monitoring in place.

That's not a model-quality problem. It's an engineering practices problem.

A survey of 650 enterprise tech leaders conducted in February–March 2026 by DigitalApplied identified five root causes accounting for 89% of scaling failures: integration complexity with legacy systems and APIs (cited by 63%), output quality degradation on edge cases (58%), absence of production monitoring infrastructure (54%), unclear organisational ownership between teams (49%), and insufficient domain training data (41%).

Read that list again. None of those are "the LLM hallucinated." They're the same things that kill any software project that ships without proper engineering discipline.

The take: Before your team commits to building an agent, ask whether you would ship any other production system without evaluation infrastructure, without monitoring, and without a clear owner. You wouldn't accept that for a payment service. An autonomous system with production access warrants at least the same standard.

The pattern that's emerging from post-mortems: teams reach for LLM capability when the real deficit is observability and ownership. The model is often the least of your problems — the surrounding engineering is.

Caveats worth keeping in mind: This data comes from VP-level decision-makers at companies with 500 to 50,000+ employees. Their definition of "production" and their risk thresholds may not match yours. Smaller teams moving faster may have different success rates. The 89% root-cause attribution covers scaling failures specifically — not all pilots that failed for other reasons.

Sources: DigitalApplied — AI Agent Scaling Gap, March 2026; Hypersense — Why AI Agents Fail in Production


2. The Claude Code Seniority Split

The Pragmatic Engineer ran a survey of 906 engineers and engineering leaders in January–February 2026. The headline result: Claude Code is the leading AI coding tool among respondents. The more interesting finding is the seniority breakdown.

Directors and senior leaders showed roughly twice the affinity for Claude Code compared to less senior engineers. Cursor adoption moves in the opposite direction — it gets less popular as seniority increases.

This is not random. Claude Code and Cursor represent two different mental models of AI assistance:

  • Cursor is a pair-programmer. It sits next to you in the IDE, auto-completes, suggests, reacts in real time. You stay in the loop at the keystroke level.
  • Claude Code is a delegate. You describe an objective, it does the work, you review the output. You're operating at the task level, not the keystroke level.

Senior engineers are used to delegating to other humans. They think in systems and objectives. Claude Code fits how they already think about getting work done. Junior engineers, still building pattern recognition and wanting tighter feedback loops, prefer Cursor's model.

The team-level implication: If your senior engineers are running Claude Code and your juniors are in Cursor, you have an invisible workflow split. This matters when they collaborate on AI-assisted work — different review assumptions, different handoff expectations, different artefacts. It's worth making explicit rather than letting it surface as friction.

Caveat: The Pragmatic Engineer survey covers newsletter subscribers — a self-selected group of experienced practitioners who care enough to read engineering newsletters. They skew more experienced than the industry at large. These numbers reflect a specific professional cohort, not a random sample of engineers. Treat the directional signal seriously; treat the exact percentages loosely.

Sources: The Pragmatic Engineer — AI Tooling 2026; Builder.io — Cursor vs Claude Code


3. The Frontier Has Moved to Inference Time

For the last several years, the competitive frontier in LLMs was training. More parameters, more data, longer runs — that's where the capability gains came from. That's no longer where the interesting action is.

The shift is to inference-time compute: spending more computation at query time — rather than at training time — to improve output quality. The core finding: a smaller model with significantly more inference compute can match a much larger model running at standard inference. Intelligence is becoming variable-cost, not fixed-cost.

What this means for engineers building on LLMs:

Cost models need to change. You can no longer think of LLM calls as fixed-cost by model. A hard reasoning task and a simple lookup task now have very different compute profiles, even on the same model. If you're not accounting for query "hardness" in your cost model, you're flying blind on your AI spend.

Latency budgets are model-dependent in a new way. Inference-time scaling approaches — chain-of-thought, tree-of-thought, repeated sampling — trade latency for quality. Whether that tradeoff is acceptable depends on your use case. For a real-time coding assistant, you may not want to spend 10 seconds thinking. For an offline audit pipeline, you might.

The infrastructure implications are significant. Inference demand in 2026 is projected to grow substantially faster than training demand — not because training is slowing, but because inference is accelerating. This is already reshaping GPU procurement and cloud pricing. If your AI budget planning treats inference as an afterthought to model licensing costs, that assumption is getting more expensive.

What this doesn't mean: Inference-time scaling doesn't fix knowledge gaps or fundamental alignment issues. A model that doesn't know something won't reason its way to knowing it. The technique is a genuine capability multiplier; it's not magic.

Sources: Sebastian Raschka — Categories of Inference-Time Scaling; LLM Inference Engineering deep dive


4. Short Takes: What Actually Mattered in March

Twelve models shipped in one week in March 2026. Most of them will matter less in six months than they seem today. Here are the three worth tracking:

GLM-5 (MIT-licensed, self-hostable, frontier-level performance). If the performance benchmarks hold at self-hosting scale — a significant if — this is the most important open-source release in months. MIT licensing with frontier capability changes the calculus for teams that can't or won't route data through closed APIs. Watch the early deployment reports carefully; benchmark performance and production performance are often different things.

Claude Opus 4.6 (1M token context window, beta). 1M context is real. For teams working with large codebases, long document corpora, or complex multi-step reasoning chains, this removes a constraint that used to require architectural workarounds. The beta caveat is real too — latency and reliability at that context length are unverified at scale. Useful to pilot now, not to depend on in production yet.

Cursor parallel subagents. This is a workflow change, not just a feature. It breaks the assumption that AI assistance is sequential. Multiple subagents running in parallel on a problem changes the time profile of AI-assisted work in meaningful ways — and introduces new questions about how you review and integrate divergent AI outputs. Worth understanding before you adopt it, not after.

The rest of March's announcements: watch the space but don't act yet.

Sources: LogRocket — AI Dev Tool Power Rankings; DigitalApplied — 12 AI Models Released in One Week


A note on sources

Everything cited above links to a primary source or a close-to-primary aggregation. Where a statistic seemed too clean — the agent pilot data especially — I've included methodology notes so you can calibrate accordingly. If you spot something that doesn't hold up, reply. That's how this gets better.


Dispatch is published every Thursday. Forward to one engineer who'd value it.

Don't miss what's next. Subscribe to Dispatch:
← Newer MCP is the integration standard. Your security posture isn't ready for it.
Powered by Buttondown, the easiest way to start and grow your newsletter.