EHO Musings logo

EHO Musings

Archives
Log in
Subscribe
July 30, 2026

Perception

It’s probably clear from the previous posting on “holding it wrong” that I hadn’t seen the Microsoft Security announcement around MAI-Flash-1-Cyber introduction into MDASH and Perception. This is what happens when you’re disconnected from the Microsoft security information firehose.🙂

The lede: Perception appears to be an impressive system for amplifying the productivity of defenders. My question is whether the industry's bottleneck has moved elsewhere, not to finding vulnerabilities, but to producing trustworthy fixes.

The main thesis of the holding it wrong post is that you have to secure the system, not pour more AI resources at the problem. Perception is an agentic system with (hopefully) good data insights and telemetry, cyber-hunting capabilities (MDASH) and an interesting set of capabilities around “blue team” and “green team”. Let’s take a closer look.

screen snip of slideware for Perception stack (from announcement video, Copyright Microsoft). Ignore the play button in this snip. 🙂

Microsoft Security Launch Event | Microsoft Community Hub

Join us to hear the latest news and announcements from Microsoft Security.‌   Speakers: Hayete Gallot Dave Weston Taesoo Kim Mustafa Suleyman

What I see in terms of the assets that Perception brings is its use of AI (agents) to increase the productivity of the various security processes (pen test, detection, vulnerability analysis, etc.) and then (I’m quoting David Weston here): “many of the folks in the SOC are not code folks so we are empowering them. We can actually write fixes with agents.” So, the SOC analyst observes an agent-coded fix and then is notified when the PR is completed and that PR was reviewed by … who?

The challenges here should be evident. Enterprise systems are already being deluged by attackers armed with working exploits, most of them discovered and authored with loosely safeguarded agentic systems and models (or white hat researchers ❤️).

What we are not lacking here is knowledge of vulnerabilities. Attackers are handing them to enterprise systems 24/7 and systems like Perception allow organizations to play “find the vuln” in a time-sensitive arms race with attackers. What we need are reliable and systematic processes to write high-quality patches and deploy them at scale. I’m pessimistic that automated code fixing, at scale, is going to be the panacea that we hope for.

The industry’s track record here is not great. The patch for Microsoft Defender CVE-2026-50656 (allowing elevation to SYSTEM) and its defense-in-depth measures allows / allowed attackers to fill the victim’s disk with large amounts of data. Patch for Windows Defender 0-day could allow attackers to fill hard disk - Ars Technica. I am fairly confident that this fix received a ton of human scrutiny, but even heavily reviewed security patches can produce unintended consequences. That should make us cautious about assuming that automatically generated patches will reliably avoid subtle (or not so subtle) regressions.

Now go back to Perception and its green team automation of fixes. No mention of which agent and model was used to code the fix. I’ve had a lot of experience over the last couple of years+ with varying quality in the code produced by agents and I read and review every line of code produced. These agents make plenty of outright dumb errors, but even scarier, subtle mistakes that are not easy to spot upon even diligent initial reading. You have an automated fix produced by an agent (again, unknown model producing the fix) and a SOC analyst not versed in the code base that he/she/they are analyzing. Are they the final human approver? At sufficient scale we all know what is going to happen: no one (human) will look at the fix.

It is undeniable we need models and agentic system to fortify the process of analyzing and defending our software systems. But the weak, soft underbelly of these systems is the production of high-quality patches which both address the underlying vulnerability and suffer a very small defect or side-effect rate. The ability of agentic coding systems to perform this task, at scale, is not present today.

There are many things to like in what Perception is attempting to do. Assimilating and analyzing telemetry (observation) is key to defense. AI-assisted triage is critically necessary, or the scale of data will overwhelm SOC analysts. But in terms of building a system that creates a high-quality and massively efficient software fix and patch pipeline? Perception appears to make substantial progress on the observation and analysis side of the problem. Whether it can substantially improve the correctness of large-scale software remediation remains an open question.

Interestingly, this reminds me a bit of Microsoft Agent 365. It represented a meaningful step forward in agentic governance by introducing structure around identity, permissions, and oversight. But it deliberately stopped short of providing runtime security infrastructure. Perception feels similar in spirit: it significantly strengthens observation, analysis, and defender productivity, yet the hardest problem remains building systems that can reliably produce, validate, and deploy high-quality fixes at enterprise scale. As with Agent 365, the enduring security boundary is not the model, it is the system that governs what the model is allowed to do.

Finally, I wanted to give a big, supportive shout-out to the Microsoft AI Red Team research engagement initiative. Yet another example of why AIRT is still the world’s best, industry-leading AI red team.
https://www.microsoft.com/en-us/security/blog/2026/07/27/enhancing-ai-security-through-global-ai-red-teaming/

EDIT: presented here without further comment: https://arstechnica.com/security/2026/07/anthropic-is-finding-bugs-faster-than-microsoft-can-fix-them/

Don't miss what's next. Subscribe to EHO Musings:
← Newer Local AI inference build Older → They're holding it wrong...
Powered by Buttondown, the easiest way to start and grow your newsletter.