The clinical AI benchmark wars — Week of July 13, 2026
Clinical AI’s benchmark war is playing out on social media. But what happens when physicians trust the wrong answer? Meanwhile, nurses say AI took their jobs, Claude joins the payer arms race, and doctors may spend more time fixing AI than it saves.
CareChronicle
Issue №06 · Week of July 13, 2026
The clinical AI benchmark wars are on. OpenEvidence demanded a retraction when it lost, then ran a study it helped design — and won. Doximity's victory lap this week comes with a footnote: a one-in-21 severe-harm rate. Meanwhile, are physicians trusting bad answers? Just remember: you're the one who's liable.
This week: New York nurses say AI took their jobs and the strikes are spreading, Claude teams up with payers through Optum and UST, and a Dartmouth study asks whether AI is actually slowing down your inbox.
The Clinical AI Benchmark Wars
Last month, a Nature Medicine study put GPT-5.2, Gemini 3.1 Pro, and Claude Opus 4.6 up against the tools built specifically for clinicians: OpenEvidence and UpToDate Expert AI. The frontier models won.
OpenEvidence demanded Nature Medicine retract the study, along with a public apology and an independent review, arguing the methods were flawed. Nature declined and pointed the company toward the standard process for filing a rebuttal.
OpenEvidence then backed a different study; they co-developed and implemented the data-collection plan, supplied the queries, and paid the physician respondents. OpenEvidence came out ahead on that one.
Doximity's head of medical AI, Louis Mullie, MD, tore it apart on LinkedIn:
"Before 'outperformed' becomes the takeaway, it's worth being clear about what this study actually is: a vendor-coordinated preference benchmark, run on vendor-filtered queries, reporting relative preference margins that say nothing about absolute clinical performance." — Louis Mullie, MD, in a LinkedIn comment
Mullie was equally unkind to OpenEvidence's new "evidence grade" badge, which collapses dozens of independently graded claims into one reassuring letter, a feature one clinical informaticist said has "the feel of trust theater."
This week, Doximity took its own victory lap. On NOHARM, a Stanford–Harvard-led safety benchmark of 1,100 physician-derived scenarios, Doximity said Ask led the tested commercial products at 75.8%, ahead of ChatGPT Health at 68.4% and OpenEvidence at 60.8%.
But dig past the LinkedIn bravado, and page eight of the study tells a different story. Doximity's own recommendations carried the potential for severe patient harm in 4.8% of cases — one in 21. Across every tool tested, severe-harm rates ranged from 2.9% to 24.6%.
These tools are already steering medical decisions. And that's a problem. Physicians show substantial automation bias toward erroneous AI recommendations, even after AI-literacy training, and keep trusting bad advice even when the patient outcome contradicts it. AI empowers smart people to confidently do stupid things.
Why are we celebrating any tool whose recommendations posed a risk of severe harm in roughly one out of every 21 cases? You're the one who's liable. Until these companies will defend your medical license, keep checking those citations.
Want earlier access and deeper analysis? Subscribe.
Mayo's AI Governance Allegations
In case you missed it: Last week, we wrote about healthcare AI moving faster than its oversight. Days later, a former Mayo Clinic research director sued the system, alleging that the team behind its MAYA assistant deleted unfavorable test results, concealed error rates as high as 67%, and retaliated when she objected; Mayo says its work complies with applicable laws and regulations and declined to discuss the case.
This week, Becker’s reports that Mayo has roughly 150 AI models deployed across the health system.
Labor Policy: The New York nurses who say AI replaced them
Replacing nurses during a nursing shortage? Montefiore laid off 12 utilization review nurses whose work was replaced by AI software, according to the New York State Nurses Association. The union says the layoffs violated safeguards negotiated after this year’s strike; Montefiore disputes that account.
The fight is spreading. California nurses protested Kaiser Permanente, and Becker’s has tracked at least 11 healthcare strikes this year in which AI governance was an issue.
Meanwhile, nursing schools turned away more than 80,000 qualified applicants because they lacked faculty. Hospitals are automating nursing roles while the system continues to restrict the supply of nurses.
This week in policy trends:
HIPAA security rules slipped again: HHS delayed the final Security Rule overhaul until at least July 2027, leaving proposed requirements such as mandatory encryption, multifactor authentication, and vulnerability scanning unresolved.
Frontier labs want regulation, largely on their terms: The leaders of Google DeepMind, OpenAI, and Anthropic now agree on independent model testing and a national governing system, though they differ over enforcement. DeepMind’s Demis Hassabis separately proposed a U.S.-led global watchdog, while OpenAI wants state laws to feed into a single federal framework.
Washington is still making the rules case by case: Recent model releases have triggered improvised negotiations and agency interventions because the U.S. lacks agreed safety standards and enough technical capacity inside government.
Payer AI: Claude, whose side are you on?
Two weeks ago, we wrote that healthcare’s AI arms race had found its battlefield: the billing office. Hospitals are aiming AI at their own charts to capture acuity; payers are aiming AI at those claims to flag and deny; hospitals are countering with agents that push appeals. AI vendors sell to every side. Patients are the ones holding the bill.
This week, Optum partnered with Anthropic to deploy Claude across its operations, offering few specifics beyond administrative support and human oversight. UST was more explicit: it embedded Claude in CarePath, a platform insurers use for claims processing, care management, and member services.
So why are denials still rising? Experian Health found that 41% of providers had at least one in ten claims denied, with denial rates increasing every year since 2022 despite an estimated $258 billion in administrative savings from automation.
Automation has shifted how reimbursement plays out, but patients are still waiting for the upside.
Everything else you need to know:
Benchmarking the process: A proposed evaluation framework for medical AI agents would score clinical reasoning, process safety, and resource use, not merely whether the final answer is correct.
Wearables without a workflow: Ninety-seven percent of surveyed physicians review consumer wearable data, but no country studied had integrated it into routine care at a rate above 6%.
Dead on arrival: Bon Secours Mercy Health’s digital chief argues that AI tools that add work to strained clinical workflows will be abandoned regardless of how well they performed in a pilot.
The AI janitor problem: A Dartmouth study of 146,000 patient messages found that physicians may spend longer correcting inaccurate or irrelevant AI-drafted replies than writing the messages themselves.
Turn it off: Health IT leaders would disable AI summaries that omit the clinical decision and alerts that turn “intelligence” into another interruption.
Governance arrived late: A Heidi Health survey found that 83% of clinicians began using AI before their employer had established a policy, guidance, or an approved tool.
Shadow health systems: A Frontiers in Public Health brief argues that tools such as ChatGPT Health increasingly shape care-seeking without being governed like healthcare delivery.
Subscribe to CareChronicle
You're reading CareChronicle. Reply with what you want more of — we read every note.