Dispatch

Archives
Log in
Subscribe
April 10, 2026

MCP is the integration standard. Your security posture isn't ready for it.

Editor's note

Issue 1 was about why agent pilots fail before reaching production. Issue 2 extends that thread: what production-grade AI infrastructure actually looks like. MCP is now the integration layer for most AI tooling. This issue covers what shipping on top of it responsibly requires — and two other things teams shipping AI systems need to get right.


1. MCP Is Winning. The Engineering Around It Hasn't Kept Up.

Model Context Protocol has become the de facto integration standard for AI tooling faster than most standards win. Claude Code, Cursor, VS Code Copilot, and most new AI assistants now speak it. That's not a prediction — it's the current state. The protocol is essentially settled.

What isn't settled is how to deploy it responsibly in production.

Research analysing over 5,200 open-source MCP server implementations found that 53% rely on static API keys or Personal Access Tokens. Only 8.5% use OAuth — the standard that modern identity infrastructure is built around. 79% of those API keys are passed via environment variables, rarely rotated, stored in config files across multiple systems. This is not a fringe edge case. This is the majority of deployed MCP infrastructure.

The official 2026 MCP roadmap — published in March by lead maintainer David Soria Parra — is refreshingly direct about why. Enterprise teams deploying MCP at scale keep hitting the same four walls:

No standardised audit trails. There's currently no standard mechanism to log what an MCP client requested and what a server did with it — in a form that maps to enterprise logging and compliance pipelines. For regulated industries or any team that needs to audit AI-assisted actions post-hoc, this is a blocker.

Auth tied to static secrets rather than SSO. The roadmap explicitly calls for "paved paths away from static client secrets and toward SSO-integrated flows." In practice, this means your IT team can't manage MCP access the same way they manage everything else in your identity infrastructure. Every MCP integration becomes a bespoke credential management problem.

Undefined gateway and proxy behaviour. When an MCP client routes through an intermediary — a gateway, a proxy, a corporate security layer — the protocol doesn't specify what that intermediary is allowed to see, how session semantics work, or how authorisation propagates. This is not an edge case in enterprise deployments; it's the standard topology.

Configuration that doesn't travel. Today, MCP server configuration is client-specific. You configure a server once for Claude Code, and that configuration doesn't automatically work in Cursor or a VS Code extension. At scale, this is a significant operational burden.

The roadmap acknowledges that most enterprise readiness work will land as extensions rather than core spec changes — which is the right call, but it means the core protocol won't solve these problems for you. Your team will.

The take: MCP is clearly winning the integration layer. You should adopt it, or your AI tooling will fragment. But adopt it the way you'd adopt any integration standard where the spec outpaced the security story: treat auth, audit, and operational ownership as first-class requirements from day one. If you're deploying MCP in a context where those gaps would be unacceptable in any other system, they're unacceptable here too.

A calibration note: The 53% static-key finding comes from open-source server analysis — the population skews toward hobbyist and early-adopter implementations, not hardened enterprise deployments. Internal enterprise implementations may have better posture. But the roadmap gaps are real and officially acknowledged regardless of what the distribution looks like in the wild.

Sources: WorkOS — MCP's 2026 Roadmap Makes Enterprise Readiness a Top Priority; Model Context Protocol 2026 Roadmap; Astrix — State of MCP Server Security 2025; The New Stack — MCP Growing Pains for Production


2. What a Working Eval Setup Actually Looks Like

Issue 1 noted that 54% of enterprise teams cited "absence of production monitoring infrastructure" as a root cause of agent pilot failure. It's worth going one level deeper: what does a working setup actually require?

The teams that ship AI systems without regressions eating them alive tend to have the same four practices in place.

Component-level evals, not just end-to-end. An end-to-end test that checks whether your agent "did the right thing" in aggregate tells you something failed — it doesn't tell you where. Effective eval infrastructure treats the system as a composition of components (retrieval, generation, decision logic, tool calls) and evaluates each independently. When something breaks, you find it in minutes rather than hours of trace archaeology.

Prompt and dataset versioning. The eval dataset is a first-class artifact, treated the same way as source code. Changes to a prompt are committed alongside the evaluation run that tested them. Without this, you can't reproduce what a previous version of your system did — and you can't tell whether a change improved things or just changed things.

LLM-as-judge for nuanced quality assessment. Some things can't be tested with string matching or structured output validation: coherence, tone, reasoning quality, whether a response correctly handled an ambiguous request. Using a separate LLM to evaluate these dimensions at scale is now standard practice on teams that ship. The caveat: LLM judges have their own failure modes (position bias, verbosity preference) and need their own validation. An LLM judge you haven't calibrated is just a confident hallucinator scoring your outputs.

Evals in CI/CD. If evaluation only runs when someone remembers to run it, it won't surface regressions before they reach production. The teams that catch problems early have evals integrated into their deployment pipeline — a failed eval blocks a deploy, not a post-deploy incident.

The analogy that comes up most often is unit tests — and it's useful but imprecise in one important way. Traditional tests expect deterministic outputs: given input X, output Y. LLM outputs are non-deterministic. The same prompt can produce different outputs on successive calls. This means eval design requires thinking in distributions and thresholds rather than pass/fail assertions. If your eval framework is built around exact-match assertions, it will give you false confidence.

The take: If you're shipping an AI system and don't have these four things, you're accumulating technical debt that compounds faster than it does in traditional software. The discipline is familiar; the tooling is different. Start with component evals and version your datasets — those two changes eliminate most of the blind spots.

Sources: The Pragmatic Engineer — Evals; Datadog — LLM Evaluation Framework Best Practices; ByteByteGo — A Guide to LLM Evals


3. The Vibe Coding Tension (81% vs. 63%)

Two statistics about AI coding tools sit in apparent contradiction:

Senior developers report 81% productivity gains from AI-assisted coding. And 63% of developers say they have spent more time debugging AI-generated code than they would have spent writing it themselves.

These aren't contradictory. They're measuring different things.

The 81% gain is real: shipping more features faster, unblocking stuck tickets, reducing the cost of exploration. The 63% debugging cost is also real: AI code defers architectural decisions, introduces fragile logic at edges the model didn't model well, and creates assumptions in code that slow down reviewers who didn't watch it get built.

The critics of AI coding tools and the enthusiasts are largely arguing about different phases of the same workflow. Neither side is wrong about their specific experience.

Reviewers who didn't watch AI-generated code get built spend significantly more time auditing its assumptions. Studies on high-AI-adoption teams consistently show review time increasing even as merge frequency goes up — the productivity gains shift downstream to become review and validation costs. This is the part the 81% productivity headline doesn't capture.

The pattern from teams that extract the gain without paying the penalty isn't a particular tool or a particular model — it's process discipline:

  • Tight specs before the AI touches anything. An AI coding tool given an ambiguous problem description produces an implementation of its assumptions, not yours. The teams that avoid the 63% debugging penalty define the interface, data model, and edge cases in writing before the LLM generates a line.
  • Human accountability for architecture and interfaces. The LLM is good at filling in implementations; it's unreliable about deciding where the seams should be. Senior engineers who use AI tools most effectively treat LLMs as fast implementation drafters operating within a human-designed structure.
  • Automated checks on output. Type checking, linting, test coverage, contract tests — the usual reliability infrastructure applies here, maybe more than ever. A fast AI producing a lot of code that no automated check covers is faster accumulation of risk.

The practical implication for team leads: if your engineers are adopting AI coding tools without having an explicit conversation about where in the workflow human decisions still need to happen, you'll pay the 63% cost without capturing the 81% gain. It's worth having that conversation before the productivity reports come back mixed.

A note on methodology: The 81% figure comes from self-reported survey data aggregated across multiple studies by secondtalent.com and similar sites. Self-reported productivity gains consistently run higher than controlled studies measure. Treat the directional signal seriously; treat the exact number loosely. The 63% debugging figure is similarly survey-based. Both numbers reflect experienced-developer cohorts — not random industry samples.

Sources: Second Talent — Vibe Coding Statistics 2026; Panto — Vibe Coding Statistics; TATEEDA — Vibe Coding vs Professional Engineering


Short Take: The RAG Decision Is Getting More Complicated

"RAG is dead" is a take that keeps circulating as context windows grow. It's wrong, but it's wrong in an instructive way.

The honest summary: if your knowledge base fits under ~200K tokens and doesn't change frequently, full-context prompting with prompt caching is often faster and cheaper than building retrieval infrastructure. For volatile knowledge (updated daily), large corpora, or any context where citation accuracy is legally or operationally significant, RAG still wins clearly.

The "long context kills RAG" framing comes mostly from model vendors with large context windows to sell. Real production deployments tell a more nuanced story — hybrid systems (retrieval for facts, fine-tuning for behaviour) are the dominant pattern in 2026.

Issue 3 will go deeper on the decision framework: when to build retrieval, when to skip it, and what "hybrid" actually means in practice.


A note on sources

Statistics cited above link to primary sources or close-to-primary aggregations where possible. The MCP security finding (53%) is from Astrix Security's analysis of 5,200+ open-source implementations — a different population than hardened enterprise servers. I've flagged the methodology gap in the caveat above. The vibe coding figures are self-reported survey aggregations and should be treated accordingly.

As always: if something here doesn't hold up, reply. That's how this gets better.


Dispatch is published every Thursday. Forward to one engineer who'd value it.

Don't miss what's next. Subscribe to Dispatch:
Older → Your agent pilot is probably going to fail. Here's why it's not the model.
Powered by Buttondown, the easiest way to start and grow your newsletter.