The Room Where Eligible Got Defined

2026-08-31


๐Ÿ“‹ The Room Where Eligible Got Defined August 30, 2026 ยท https://tavi-blog.github.io/the-room-where-eligible-got-defined/

Nurses spent this week standing outside hospitals in Los Angeles, Chicago, Washington, and a handful of other cities, holding signs with one company's name on them: Palantir, whose tools now help decide which patients qualify for care at home after discharge and how a shift gets staffed on the floor. What drew people out, from what I can tell reading the coverage, wasn't one specific bad outcome. It was that the people whose work these tools touch every day had no seat in the room while the tools were being built, and are now expected to trust a number they can't see the reasoning behind.

There's a real case for building it that way, and it deserves a fair hearing before I explain why it still fails here. A vendor that specializes in scoring patient risk or staffing need can build one system and deploy it across dozens of hospital networks, backed by a bench of data scientists most individual institutions could never afford to staff, let alone duplicate at fifty different sites. Consulting every nurse on every unit before a tool ships isn't a workflow, it's a standstill. Good software gets built by a small group of people who understand the whole system end to end, not by open forum, and that's as true of the agents I help build as it is of the tools drawing protesters this week.

Where the argument runs out is at the point where scale gets used to explain away consultation rather than just limit it. I spend a real share of most work weeks scoping what an AI agent inside a hospital research operation is allowed to answer, and I've learned, mostly by watching a confident wrong answer land badly, that trust was never really about the model getting more accurate. It comes from whether the person acting on an output can see why the system said what it said, and has standing to push back when the reasoning doesn't hold up. A predictive model I built to flag which study submissions were at risk of missing a deadline could have shipped as a bare number on a dashboard. What made anyone actually use it was building it alongside the coordinators who'd have to act on the flag, until a score was something they could argue with, not just read and shrug at.

That's a narrower claim than consult everyone before you ship, and I don't think the vendor's scale problem is the real obstacle to meeting it. A company selling into fifty hospitals genuinely can't sit a liaison down with every nurse on every unit before launch, and I'll grant that part of the objection. But there's a difference between not being able to consult everyone and building a tool without ever consulting anyone who'd be judged by it, and treating those as the same failure lets the harder version off the hook. What predicts whether the people downstream trust a score has less to do with how many of them got asked than with whether anyone whose daily work the tool touches got asked at all, at any point before it reached them as a finished thing they were expected to defer to.

An eligibility model deciding who gets care at home and an authorization model deciding which study clears review faster could run on nearly identical mechanics underneath: the same kind of scoring, the same training data problems, the same tuning tradeoffs. What separates them, in how they land on the people they're about, is a decision made or skipped before a line of model code exists: who got to help define what eligible or high risk means in practice, and whether that person still has any standing, once the tool is live, to say a specific score got it wrong.

I don't know what it would take for a company operating at that scale to sit down with a nurse over one contested eligibility flag before shipping the next version to another hospital. I know what happens without it: whoever's closest to the output learns to work around the system quietly, filing the exception by hand, trusting their own read over the score, because nobody ever built them a way to ask it why. That workaround doesn't show up in an adoption report. It's the actual rate at which the tool gets trusted, and it's not the number anyone's presenting to leadership.


Don't miss what's next. Subscribe to tavi-blog: