The Signal - daily AI evolution  logo

The Signal - daily AI evolution

Archives
Log in
Subscribe
August 25, 2026

The Signal: Bio/Health — Week of Aug 24, 2026 — Edition #2

The Signal: Bio/Health — Week of Aug 24, 2026 — Edition #2

The Read

The week's headline is not a molecule; it is a hit rate. Anthropic published a protein-design campaign in which its models designed binders against fifteen targets, had them physically synthesized and measured by two independent labs, and hit fourteen — 354 confirmed binders out of 1,320 designs, at roughly twice the industry baseline. Every prior generative-biology milestone we have logged sat at the same rung, but this one arrives with third-party synthesis, affinity measurements, and a target it failed on, which is exactly the shape of a result you can underwrite. Meanwhile the regulatory conversation stopped being about whether a doctor reviews the AI and started being about whether one should have to: a JAMA viewpoint argued that mandated human oversight will soon make care worse, the AMA published a framework insisting the physician's role is durable, and the FDA cleared a robot that draws blood with nobody's hands on it. The design layer got measurably better this week. The oversight layer got measurably more contested. Those are the two stories, and they are not the same story.

Tide status

No tide movement this week. Both standing tides hold, one with a sharper evidence base.

Biology is becoming an engineering discipline — holds. The Anthropic/Adaptyv result replaces Arc's Evo-generated phages as our best generative data point, but it does not climb the ladder: both sit at [wet-lab]. What changed is the quality of the rung, not the height. Phages were a striking demonstration with in-house validation; binders with third-party SPR curves, disclosed failures, and quantified affinity deltas are the kind of evidence a board can act on. The tide moves when a computationally designed asset produces a human outcome — INTerpath-001's presented Phase 3 data, whenever it lands, remains the nearest test.

The binding constraint is disease understanding, not molecular design — holds, and this week supplied an unusually precise restatement of it. If protein design is now running at 2x baseline hit rates, the bottleneck has moved further downstream, and a16z's "Oracle Problem" essay names where: we do not know how to tell whether a clinical AI was right. Abundance at the design bench does not fix scarcity at the ground truth.

Waves

1. Generative design crossing into wet-lab function — MOVED [wet-lab]

Anthropic's campaign is the first time a frontier lab has run generative protein design end-to-end and had a third party check the work at scale. Adaptyv Bio and Twist Bioscience synthesized and characterized the designs: 95% expression, SPR at five concentrations in duplicate, best 15-PGDH binder at 33.4 nM against a prior competition winner's 1.7 µM, RBX1 at 3.9 nM against 25.7 nM, TREM2 at 80% success against 38.3%. One target (MBP) produced nothing; one (GDF-8) was excluded for measurement quality. Anthropic's own caveat is the honest one — minibinders are not a standard therapeutic modality, and binding is the first step of drug development, not a proxy for it. The campaign also needed specialized folding models and GPU access that no ordinary Claude user has, so this is a lab capability, not a product.

Roadmap implication: the design layer is commoditizing faster than the validation layer can absorb it. Underwrite wet-lab throughput, characterization capacity, and the CRO/foundry tier — Adaptyv is as much the story here as Anthropic. The scarce asset in 2027 is not a model that proposes binders; it is the bench that can tell you which ones are real. Anthropic · Adaptyv Bio

2. Clinical-AI autonomy vs. mandated human oversight — NEW [policy]

Ezekiel Emanuel and Neal Khosla argued in JAMA that regulators should not require a human in the loop, on the grounds that autonomous AI will shortly outperform doctor-plus-AI teams. Their evidence is real but uniformly synthetic: ChatGPT o3 diagnosing correctly first in 60% of 377 cases against 15.9% for internists, GPT-4 alone at 92% on diagnostic reasoning against 76% for physicians given access to it, and a meta-analysis of 106 experiments finding humans degrade AI performance where the AI is superior. Two days later the AMA and the Digital Medicine Society published a framework asserting five enduring physician responsibilities that no amount of capability displaces. Disclosure matters here: Khosla runs Curai Health, an AI telemedicine company with OpenAI ties.

Roadmap implication: the human-in-the-loop requirement is the single largest hidden variable in clinical-AI unit economics. Every business model in the category prices supervision in or out. Watch whether the FDA's forthcoming guidance treats oversight as a safety control or a legacy assumption — and note that nearly all evidence cited against oversight comes from simulations, not delivered care. We would not act on it yet. JAMA viewpoint · AMA/DiMe framework

3. The regulatory regime for adaptive systems — MOVED [policy]

Rick Abramson, who runs the FDA's Digital Health Center of Excellence, told STAT the agency will issue formal guidance on generative-AI-enabled devices, plus specialty guidance on particular topics. The Aug 18 discussion paper floats a "competency-based approach" borrowed loosely from how physicians are assessed, and — more consequentially — asks whether the agency should accept greater premarket uncertainty in exchange for stronger postmarket monitoring. Experts cautioned against reading the competency framing too literally; it remains a device framework underneath. The catch is legal: the FDA may need new statutory authority from Congress for some of the postmarket concepts it is contemplating, which converts a guidance timeline into a legislative one.

Roadmap implication: the comment docket (FDA-2026-N-7874) closes Oct 19. That deadline is now the highest-leverage cheap action available to anyone with a position on this. Our standing action item stands. FDA discussion paper · STAT

4. Abundance of molecules vs. scarcity of validated mechanisms — the measurement problem moves up the stack

a16z Bio+Health's "Oracle Problem" essay makes an argument we have not seen stated this cleanly: healthcare AI cannot be evaluated because there is no ground truth to evaluate against. Physician identity explains between 7% and 77% of the variation in treatment choice across fifteen paired operations — for partial versus total knee replacement, patient characteristics, comorbidities, facility and year explain 3.4% of the variance; adding the surgeon's identity takes it to 14.8%. A model can disagree with the recorded decision and still be clinically defensible. Worse, five of six public healthcare AI benchmarks give models less context than the median patient record (~8,500 tokens), several under 200 tokens. In a 19-way diagnosis task, permuting answer order changed model answers; adding one sentence defining a threshold flipped which model won. Owl Posting's organoid survey lands adjacent: across chemotherapy response, cystic fibrosis, Zika and immune organoids, the 3D structure was strictly necessary in roughly one case out of four.

Roadmap implication: benchmark scores in clinical AI are close to uninformative as procurement inputs, and vendors know it. Ask what a system predicts, over what horizon, against what outcome that was measured after the fact — and treat any leaderboard claim without a stated grader, prompt and context window as marketing. a16z Bio+Health · Owl Posting

Ripples

1. Anthropic hits 14 of 15 protein targets, with outside labs holding the pipette [wet-lab]

354 confirmed binders from 1,320 designs; hit rates of 22.6% (Opus 4.8, multi-target, 48h) to 35.1% (Mythos Preview, single-target, 24h) against a 10–15% industry baseline. Separately, Claude processed NMR and LC-MS datasets in 23 and 19 minutes against a lab's roughly four days to a finished report. The daily Signal ran this Aug 18; the reason it belongs here too is the rung, not the headline — this is validated bench chemistry, not a benchmark.

So what: the fastest-improving input to drug discovery is now the cheapest one, which means the returns keep migrating to whoever owns validation and clinical assets. That is the same conclusion Pathos AI reached last week by spending $2.2B on clinical-stage chemistry. Anthropic

2. FDA authorizes the first fully autonomous blood-draw robot [approved]

Vitestro's Aletta cleared De Novo on Aug 19. It positions the patient's arm, finds a vein using near-infrared imaging and Doppler ultrasound to distinguish it from an artery, applies the tourniquet, preps the skin, inserts and disposes the needle, swaps tubes and applies the bandage. Trials showed draw success comparable to or better than trained phlebotomists across difficult-access patients and varying skin tones; adverse events uncommon and mild. One phlebotomist can supervise three units.

So what: this is the autonomy debate escaping the diagnostic screen and entering the procedure room, and the FDA authorized it the day after publishing a discussion paper about how carefully to supervise generative systems. Note the supervision ratio — 3:1 is the actual economic claim, and it prices labor, not accuracy. FDA

3. Tempus clears an AI that reads pulmonary hypertension off a standard ECG [approved]

Cleared Aug 24, Tempus ECG-PH flags patients with mean pulmonary artery pressure above 20 mmHg from a routine 12-lead, indicated for adults 40+ or younger adults with cardiovascular symptoms. It is Tempus's third cleared cardiovascular algorithm, after atrial fibrillation and low ejection fraction. No sensitivity, specificity, AUC or validation cohort size has been published.

So what: the commercially live version of medical forecasting is not a foundation model, it is a cheap existing signal read for a condition it was never ordered to detect. Also note the disclosure gap — a clearance with no published performance data is a reimbursement conversation waiting to happen, and the CMS NTAP cycle is where that conversation gets settled. Cardiovascular Business

4. FDA commits to generative-AI guidance — and may need Congress [policy]

Abramson's confirmation to STAT that guidance is coming, with specialty guidance behind it, is the first firm signal since the discussion paper. The postmarket-monitoring-for-premarket-uncertainty trade is the live idea, and parts of it may exceed the agency's current authority.

So what: a guidance timeline you can plan around becomes a statutory timeline you cannot. Anyone building an adaptive clinical product should assume the premarket bar moves slowly and the postmarket obligation arrives first.

5. Arc Institute's 2026 Virtual Cell Challenge goes zero-shot [in-silico]

Opened Aug 20: predict CRISPRi knockdown responses in six cell lines the model has never seen perturbed, with no training set, scored on an aggregate of six metrics rather than one. $100k grand prize, test set Oct 22, submissions close Nov 5; NVIDIA, 10x Genomics and Ultima Genomics sponsoring. The 2025 edition drew 5,000+ registrants from 114 countries and 1,200+ teams.

So what: the interesting design choice is the six-metric aggregate — a direct response to single-metric gaming, and the same instinct the Oracle Problem essay applies to clinical AI. Watch the zero-shot scores in November; cross-cell-type generalization is the capability that would make virtual cells useful rather than impressive. Arc Institute

6. Evaxion scraps its solid-tumor vaccine work to concentrate on an AI-designed glioblastoma candidate [preclinical — unverified]

Announced Aug 18, refocusing on the Duke-partnered program targeting endogenous retrovirus antigens.

So what: a small company narrowing to one AI-designed asset in the hardest solid tumor there is reads as conviction or as cash discipline, and from outside you cannot tell which. We still cannot confirm the program's exact clinical stage from public filings — which is itself the finding, and why it stays flagged for verification. Fierce Biotech

Pipeline watch

The scoreboard for whether AI drug discovery is actually working. Movement this week is at the platform end, not the clinic.

Company Asset Stage Change this week
Merck / Moderna intismeran autogene (INTerpath-001) [Phase 3] Added. Phase 3 melanoma success announced Aug 19; up to 34 algorithmically selected neoantigens per patient. Zero effect sizes disclosed. Data expected at a medical meeting — ESMO, Madrid, Oct 23–27, is the likely venue. Honest caveat: the neoantigen selection predates modern AI; this validates computational personalization, not deep learning.
Insilico Medicine rentosertib (ISM001-055), IPF [Phase 3] No change. NCT07687459, 320 patients, 47 China centers, 52-week FVC endpoint. Still the category's marquee test article.
Recursion REC-617 (CDK7) and portfolio [Phase 1] No change. TUPELO data for REC-4881 due Nov 2026.
Isomorphic Labs first candidates preclinical No change — still no disclosed clinical candidate, nothing since Jul 16. Demis Hassabis's move to an Alphabet chairman/Chief Scientist role explicitly names Isomorphic as a focus; read that as attention, not as an IND.
Anthropic / Adaptyv Bio minibinder platform [wet-lab], no clinical asset Added. Tool layer, tracked for downstream partner assets.
Evaxion AI-designed GBM vaccine (Duke) early — unverified Refocused Aug 18; stage still unconfirmed.
insitro / Xaira / EvolutionaryScale / Chai platforms preclinical / no clinical asset No change; quiet week across all four.

Zero full FDA approvals of an AI-discovered drug to date. That number is the whole scoreboard, and it did not move.


The Signal — Bio/Health, from The Excelsior Group. https://excelsiorgroup.ai/insights/signal/bio/

Don't miss what's next. Subscribe to The Signal - daily AI evolution :
Older → The Signal: Human Advancement — Week of August 24, 2026 — Edition #2
Powered by Buttondown, the easiest way to start and grow your newsletter.