2026-08-07
🧪 A Testing Environment Is Not a Checkbox August 6, 2026 · https://tavi-blog.github.io/a-testing-environment-is-not-a-checkbox/
A survey of health system leaders published this week asked a fairly plain question: before an AI tool goes anywhere near a real workflow, does the organization actually have a place to test it. The answer, from research out of the Center for Connected Medicine and KLAS, was that most say yes and mean something much thinner than what the word implies. Health systems report validating their AI tools before deployment far more often than they have the infrastructure to do that validation thoroughly, and almost two-thirds say they don't have anything resembling an advanced AI strategy to govern the decision either way. Adoption hasn't slowed down waiting for any of that to catch up. Ambient documentation tools lead the list of what's already live, followed by revenue cycle automation, imaging support, and decision tools embedded directly in the record.
The instinct to read that gap as recklessness is fair, and it deserves a real hearing before I complicate it. Waiting for a fully built governance framework before touching a live workflow sounds responsible right up until you notice what the wait actually costs: a clinician still writing the note by hand, a coding backlog that keeps growing, a queue that could have been triaged faster months ago. An institution that refuses to move until every testing environment is formally provisioned isn't avoiding failure, it's choosing a slower, quieter version of it. Unaccountable is still a way of failing the people the tool was supposed to help.
But I've built the thing this survey is describing, at a scale small enough that nobody would have called it AI strategy, and the gap it names is one I recognize specifically rather than abstractly. A predictive model I built for study approval timelines had nowhere sanctioned to be tested before it touched anything live. There was no staging environment waiting for it, no existing protocol for how long to run it alongside the old way of estimating a timeline before trusting its output over that, no institutional definition of how wrong a prediction had to be before it should get pulled. I had to invent all of that myself, because the model wasn't the kind of tool anyone at the institution had a template for validating. It didn't diagnose anything. It didn't touch a patient. It predicted how long a piece of paperwork would take to clear, which turns out to be exactly the category this report is describing when it says systems are moving faster than their own governance.
What the survey calls a testing environment and what I actually had were two different things wearing the same word. A testing environment, in the way a governance framework means it, is infrastructure: a sandboxed replica of production, a formal sign-off before promotion, a person whose job it is to say yes or not yet. What I had was a personal decision to keep checking my own model against real outcomes for weeks longer than anyone required, because I didn't trust it yet and no one else was positioned to tell me when I should. That's not nothing. It's also not something you could point to on an org chart, and it isn't the kind of thing a survey of health system leaders would ever pick up as an answer, because nobody asks "does an infrastructure gap like this get quietly covered by an individual" as a checkbox on a KLAS questionnaire.
That's the part of this report I keep sitting with longer than the topline number. The gap between adoption and governance isn't only a resourcing problem that a bigger budget line eventually closes. Some of what's currently filling it is individual discipline, happening in places nobody audits, and it holds only as long as the person doing it feels like keeping it up on their own time. A framework that formalizes validation doesn't just catch the tools that would have failed. It stops the whole system depending on whether the person building the next one happens to be the kind who keeps checking their own work after everyone else has stopped asking.
Don't miss what's next. Subscribe to tavi-blog: