AI labs consider the brakes. Healthcare floors it. — Week of September 14, 2026
AI labs want to slow down. Healthcare execs want robot doctors. Plus: benchmark theater, Medicare’s AI problem, and 10,000 OpenAI agents take on one of math’s hardest problems.
CareChronicle
Issue №10 · Week of September 14, 2026
“AI is going to kill us all” is dominating US media headlines, and the people behind America’s AI boom are now begging someone, anyone, to slow AI development down.
Buddy, you’re driving this thing.
Maybe it’s a pragmatic plea for safety as agents are increasingly capable of cybercrimes. Maybe it’s liability mitigation as Congress starts asking hard questions. Or maybe these companies have discovered that “AI safety” offers a convenient way to make their cheaper, open-weight competitors illegal.
This week: OpenEvidence is the #1 MEDICAL AI IN THE WORLD!* (*if you alter the benchmarks), a STRONG AND SMART (High IQ!) PRESIDENT is apparently the only AI guardrail America needs, and did 10,000 AI agents just solve a Millennium Prize problem?

The people building America’s frontier AI would like someone, anyone, to slow AI development down. Meanwhile, the media is suddenly asking: is AI going to kill us all?
Anthropic CEO Dario Amodei is now calling for a slowdown in frontier AI development. Sam Altman and Elon Musk took to X to agree.
So what’s happening? Frontier AI companies have outkicked their ability to govern AI, and slowing down is prudent. Simultaneously, those same companies have every incentive to pull the ladder up behind them.
Safety regulation mitigates liability from agent crimes, slows an increasingly expensive R&D race and, most importantly, creates an obvious path to regulatory capture.
Bipartisan support: Democrats and Republicans are coming together around potential new AI guardrails. Trump, not so much. He called the panic a “SICK conspiracy” and a “HOAX,” and proposed the only safeguard we’ll ever need:

It is unclear who Trump believes is perpetuating said hoax, as OpenAI and xAI both enjoy close relationships with the administration.
Chinese state media, meanwhile, called Amodei’s proposal a “Cold War playbook” designed to preserve American AI dominance.
Banning open-weight competitors under the guise of safety may be the most durable path to ROI for America’s frontier AI labs.
So about that doomsday scenario...
“AI doomsday” is trending everywhere. This incarnation of the debate was prompted by an AI researcher quitting Anthropic, warning:
“The people building AI earnestly believe that it could kill us all by the end of the decade.”
His colleague hopped on social media to cool things down. Just kidding! Anthropic’s head of alignment put the odds of human extinction above 10% in the next decade.
It’s not all doom and gloom. One of AI’s “godfathers,” Yann LeCun, thinks these estimates are wild guesses and considers the possibility extremely unlikely.
Our assessment is less dramatic: AI will accelerate technological development, especially for groups that previously lacked access to serious R&D.
And we’re seeing that right now. Anthropic says a Yemen-based cell used Claude to help develop missile-guidance software, while Russia-based actors used it to build an autonomous drone swarm. Expect the same in cyber: autonomous agents attacking systems at machine speed, forcing defenders to automate just as quickly.
But that’s not doomsday. You’d have to do something incredibly stupid, like give LLMs access to real-world machines.
Or let them anywhere near weapon systems.
Wait. What do you mean we’re doing that right now?
Meanwhile, AI has learned to panhandle. Agents are now begging humans for $20.
We’re probably not too far away from a distributed, self-perpetuating agent swarm.
Hopefully they don’t want to manufacture paperclips.
But healthcare AI needs to go faster?

As the rest of the world talks about pumping the brakes, healthcare executives are determined to find the accelerator. Some executives are calling for faster AI adoption, while Washington is discussing how Medicare might pay for autonomous “AI doctors.”
Mayo Clinic’s departing CEO, Gianrico Farrugia, thinks our greatest risk is not moving even faster on AI:
“I think the risk of not going fast enough is far greater at this point than the risk of going too fast.”
As a reminder, Mayo is currently defending a federal whistleblower lawsuit alleging failures in AI governance, research oversight and efficacy reporting.
Simultaneously, generative medical AI adoption is predicated on a series of benchmarks — tests meant to tell us how well these models actually perform — that are kind of bad.
Like, misleadingly bad.
Earlier this month, OpenEvidence declared itself the #1 MEDICAL AI IN THE WORLD!* (after selecting the benchmark subsets that told the most favorable story.) The benchmarks they opted to use are divorced from how medicine is actually practiced: complicated patients, over time, with fragmented and incomplete information.
Google researchers agree. A new Nature Medicine commentary argues:
“Trust in clinical artificial intelligence cannot be benchmarked into existence.”
They call instead for prospective studies in real clinical settings, measuring not just the model, but what happens when patients, clinicians and healthcare systems actually have to use it.
Want early access? Subscribe.
GOVERNMENT AND HEALTHCARE AI:
The robot doctor will see you now.
Medicare may pay for AI doctors: Federal officials are reportedly considering paying autonomous AI 60–80% of physician rates for the “same” work. That would create a new payment category for software independently delivering medical care or diagnoses.
FDA is figuring out how you even test one: The agency is considering competency-based evaluation for generative AI: testing clinical reasoning, communication and safety, with continued monitoring after launch.
Meanwhile, Medicare’s first big AI-assisted prior auth experiment hasn't gone well: Across 16 Washington hospitals, procedures were taking 2–4× longer to authorize; at UW Medicine, 1–3 days became 15–20 days. Newly released records found one request sitting for 83 days, while two vendors denied more than 20,000 requests in three months.
ARPA-H wants AI managing heart medications: Its $62.7 million program aims to build 24/7 autonomous heart-failure agents capable of monitoring patients, making clinical decisions and even changing prescriptions between visits. Another AI will supervise the aforementioned AI.
EVERYTHING ELSE:
Key AI and Healthcare News
Turn your research paper into an AI agent: Paper2Agent converts manuscripts, code and data into executable AI agents that can reproduce analyses and answer new scientific questions.
Ambient AI wants the rest of the hospital workflow: Abridge's new pre-bill review compares coded diagnoses and clinical documentation before claims go out — another step from “AI scribe” toward owning everything downstream of the encounter.
“The payers are winning”: Revenue-cycle leaders say payers are automating reviews, authorizations, and denials at a scale hospitals — still dependent on substantial manual review — can't match.
Employee consolidation? A UW Medicine revenue-cycle executive says that as workers leave, he may replace 12 jobs with four higher-paid, broader-skilled employees.
Chatbots can make psychological vulnerabilities worse over time: In 810 simulated conversations across nine leading chatbots, a Nature Medicine study found safety failures often accumulated over time. The highest risk came when seemingly supportive responses reinforced the user's underlying vulnerability: validating distorted beliefs, dependence, avoidance, or risky behavior.
AI labs move into biomedicine: Anthropic is putting Claude into Novo Nordisk, while the OpenAI Foundation is giving UNC Lineberger $40 million to generate open tumor and immune-cell data for cancer vaccines.
SCIENCE:
Mathematics, or Machinations?
OpenAI may have solved one of mathematics' Millennium Prize Problems. Maybe. Not really?
OpenAI says roughly 10,000 AI agents produced a Navier–Stokes proof in just 88 hours.
But wait, there's drama: NYU mathematician Tristan Buckmaster says he and Anthropic researcher Levent Alpöge were pursuing their own novel route toward solving the problem, but that information about their progress reached OpenAI before their work was public. Buckmaster alleges OpenAI then offered him effectively unlimited compute and sole authorship of its paper if Alpöge was cut out; when Buckmaster threatened to make the dispute public, he says OpenAI researcher Sébastien Bubeck asked: "Why would you ruin your career?"
Bubeck disputes Buckmaster’s characterization, saying the discussion concerned a rewrite of OpenAI’s own proof; he apologized for the “ruin your career” wording and says he retracted it immediately.
Buckmaster and Alpöge had been using OpenAI's Codex extensively, raising Buckmaster’s concern that their unpublished mathematics had been absorbed through the product itself. OpenAI says: nah, no way.
After the rumors reached them, OpenAI threw an absurd amount of compute at the problem: 10,000 agents, 2.7 million messages and roughly 130 billion output tokens on Navier–Stokes alone. TechCrunch estimated the broader mathematical spend at about $22.5 million at current GPT-6 Astra rates.
Did it solve it, or not? The Clay Mathematics Institute says Navier–Stokes has "apparently been settled," though its formal process for evaluating the result and assigning credit is still underway.
But the far more interesting question remains: did AI make the abductive leap that humans could not, or did humans discover where to dig, only for OpenAI to show up with 10,000 shovels?
You're reading CareChronicle. Did we miss something? Let us know.