2026-07-19
๐๏ธ Nobody Hacked the Model July 18, 2026 ยท https://tavi-blog.github.io/nobody-hacked-the-model/
A phishing email got answered in January, and by the time the healthcare company that got compromised finished notifying anyone, it was June. Xsolis, the utilization review vendor whose scoring feeds into workflows at Humana, Mayo Clinic, and a handful of other systems, disclosed a breach touching 1.4 million patients five months after it happened. Names, birth dates, social security numbers, insurance details, treatment history. The kind of record that gets described in coverage as "an AI company's dataset," as if the AI part is where the interesting failure lives.
It isn't. I spend a good part of most work weeks doing something close to the opposite of what that phrase implies: taking data that lives in five or six systems never designed to talk to each other and forcing it into something a dashboard, or a model, can actually read. Ambiguous join keys. Duplicate values sitting three layers deep in a lookup column that only surface once you build the relationship and watch the numbers come out wrong. None of that is the part of AI anyone finds interesting. It's the part that has to happen correctly, quietly, thousands of times, before anything interesting can happen at all.
There's a real case for handing that work to a specialized vendor instead of building it in house. A hospital's IT function is stretched across clinical systems, physical infrastructure, and a hundred point solutions that predate anyone currently on staff. A company whose entire business is "we process this category of data securely" should, in theory, be better at the securely part than an institution doing it as one responsibility among fifty. Specialization is supposed to raise the floor. That's not a naive argument. It's the same logic that justifies most of the vendor sprawl inside any large hospital network, including the one I work in.
What nobody says out loud when a system hands data processing to a vendor is that you're not only buying their security expertise, you're buying your own blindness to the risk until they decide it's time to tell you. Rochester Regional Health found out 18,600 of its patients were exposed because Xsolis told them, on Xsolis's schedule, five months after Xsolis found out itself. The hospital didn't fail to secure that data. It never had the ability to check. That gap between "responsible for the patient" and "able to see what's happening to their record" is the actual structure being described every time a story like this runs, and it rarely makes it into the second paragraph.
Every dataset that eventually gets called training data or a model input spent most of its life as somebody's ongoing headache: a spreadsheet a coordinator half trusts, a form field that three different teams populate three different ways, a validation flow wrapped in error handling because a silent failure is worse than one that makes noise. The AI framing skips past all of it, because "a vendor got phished" reads as a smaller sentence than "a machine learning company's dataset was breached," even though both sentences describe the exact same five months and the exact same spreadsheet-adjacent mess underneath them.
I don't think the fix is more scrutiny of the models these companies eventually ship. The model was never the weak point in this story, and it almost never is. The weak point was a person opening an email, inside a system built to move sensitive records between institutions faster than any of those institutions can independently verify what's happening to them once they leave the building. That's not a problem AI governance is built to catch, because it isn't really about the AI. It's about how many hands healthcare data passes through, and how few of those hands belong to anyone with the authority to say no.
The patients who got a letter this month don't know or care whether the failure happened inside a model, an inbox, or a spreadsheet someone forwarded without thinking twice. They know a company they never signed up for had their records, and that the hospital they trusted found out about it the same way they did: by being told, months late, by someone else. I keep thinking about how much of my own job is trying to make sure that gap never opens on our end, one join key and one validation check at a time, and how little that kind of work shows up in anyone's account of what AI in healthcare is actually made of.
Don't miss what's next. Subscribe to tavi-blog: