The Signal - daily AI evolution  logo

The Signal - daily AI evolution

Archives
Log in
Subscribe
September 1, 2026

The Signal: Bio/Health — Week of Aug 31, 2026 — Edition #3

The Signal: Bio/Health — Week of Aug 31, 2026 — Edition #3

The Read

Last week a frontier lab reported a 14-of-15 hit rate on designed protein binders and we called it the best generative-biology evidence on the board. This week the field got its first blinded, prospective, independently scored report card on AI molecular design — and the grade is a solid, unglamorous incomplete. In the AIntibody benchmark, twenty-nine organizations submitted 511 antibodies against a common target with the answers hidden; the best AI-matured antibody came in at 95 pM against 113 pM for three months of ordinary phage display, a difference the authors call statistically indistinguishable. In a second challenge, models beat simply picking the most abundant clone less than 14% of the time, while picking at random beat them 39% of the time. That is what a real benchmark does to a field that has been grading its own homework. And on the same day the paper's press cycle ran, one participant issued a release headlined "AI Design Surpasses the Best Experimental Result." Both statements describe the same dataset. Only one of them survives reading it.

Meanwhile the money told a different and more interesting story than the models did. Insilico Medicine posted $106.3M of first-half revenue and $35.5M of net profit — the first genuinely profitable half from a listed AI-drug-discovery company — and its founder spent the weekend publishing an essay arguing that the private comparables in his own sector are naked emperors. When the one company with audited numbers is the one calling the bubble, that is worth more than a leaderboard.

Tide status

No tide movement this week. Both standing tides hold; one of them got its first adversarial test and survived with a smaller claim.

Biology is becoming an engineering discipline — holds, with a narrower rung. AIntibody is the first time generative molecular design has been scored blind, prospectively, against experimental ground truth it could not see, by people with no stake in the answer. The result is not that AI design fails; it is that AI design works inside defined optimization settings and generalizes poorly outside them, and that no single algorithmic approach dominated. That is precisely the shape of a technology at [wet-lab] rather than one climbing the ladder. Nothing here contradicts last week's binder campaign. It does mean the honest version of the claim is "competitive with three months of conventional work, sometimes, on tasks you have specified carefully" — not "superhuman."

The binding constraint is disease understanding, not molecular design — holds, and got two independent restatements. Daphne Koller, in a long STAT interview, put the bottleneck at the bridge between a cellular mechanism you can measure in a dish and a clinical outcome you can defend: "the mapping, the bridging between those, is where drugs often fail." And a Cell perspective from Andrea Califano and colleagues argues the molecular-scale wins — folding, mutation effects — do not transfer up to cellular and multicellular biology, and proposes fifteen challenges as the field's agenda. Two very different rooms, the same finding.

Waves

1. Generative design crossing into wet-lab function — MOVED (first independent scoring) [wet-lab]

AIntibody ran three blinded challenges against the SARS-CoV-2 RBD: affinity maturation, cluster ranking, and de novo CDR design. Challenge 1 drew 165 submissions from 25 organizations; only 13% achieved a 20-fold or better improvement, and 63% produced developable binders. The best AI entry (95 pM) and the best experimental entry (113 pM) were statistically indistinguishable, and a purely statistical consensus baseline placed third at 540 pM. Challenge 2 is the uncomfortable one: models beat picking the single most abundant clone only 9.8–13.8% of the time, while random selection improved on it 39% of the time. Challenge 3's top affinity result — 2.9 pM — failed developability on a HIC column, and 30.4% of submissions did not bind at all.

Roadmap implication: the design layer is real and it is narrow. Underwrite the benchmark infrastructure and the wet-lab capacity that scores it, not the leaderboard position — this week the same dataset produced both "AI performs in defined optimization settings and generalizes poorly" and a vendor headline claiming AI surpassed the best experimental result. Any diligence process that cannot tell those apart will overpay. Ask for the blinded result, the developability data, and the failures. Nature Biotechnology · critique

2. The AI-drug-discovery show-me window — MOVED (first hard financials) [deal]

Insilico Medicine reported first-half 2026 revenue of $106.3M, up 287.2% year over year, at a 90.3% gross margin, with $35.5M of net profit and $584.8M of cash and investments. Drug discovery and pipeline revenue was $103.1M against $2.7M of software — the platform sells programs, not seats. Days later, founder Alex Zhavoronkov published an essay arguing that private AI-biotech valuations have detached from fundamentals, citing a platform company raising at $2B against a sub-$1B addressable market and a company with no clinical assets or disclosed targets above $15B. His argument, stripped of the sales pitch attached to it: structure prediction is one skill family out of roughly 1,200 required to produce a medicine, and capital is flowing to the skills that are easiest to demonstrate rather than the ones closest to the decisions that kill programs.

Roadmap implication: he is talking his book — he runs the listed comparable that looks cheap if he is right, and his firm sells the benchmark suite he proposes as the remedy. Discount accordingly. But the underlying test he proposes is the correct one and it is cheap to apply: ask what a computational skill does to the probability a safe and effective drug reaches a patient, then ask what that drug does to global health. Both answers are usually small, and neither appears on a leaderboard. The scoreboard number itself did not move — still zero full FDA approvals of an AI-discovered drug. Insilico H1 results · Zhavoronkov essay

3. Abundance of molecules vs. scarcity of validated mechanisms — MOVED (the measurement problem gets a research agenda)

Koller's STAT interview is the clearest public statement yet of what insitro thinks it is doing differently: combining human data (omics, pathology, radiology, outcomes) with perturbational lab data to establish causal links between molecular changes and clinical manifestations, rather than optimizing a single modality. Her worked example is instructive — LDL and HDL are both correlated with heart disease, only LDL is causal, and the drugs that raised HDL failed. She also flags AI-discovered biomarkers as the realistic lever on trial duration, on the HIV viral-load precedent, but insists they need a formal regulatory framework proving causal relation before they can shorten anything. The Cell perspective arrives at the same wall from the academic side, and AIntibody's Challenge 2 supplies the empirical version: if a model cannot beat clone abundance at ranking, it has not learned the biology, it has learned the dataset.

Roadmap implication: the near-term commercial value in AI×bio is more likely to sit in validated surrogate endpoints than in molecule generation, because endpoints compress the expensive part of the timeline and molecules do not. Watch which companies are building the regulatory case for a new biomarker, not which are publishing hit rates. STAT AI Prognosis · Koller, "Drug Discovery Has No Magic Wands" · Cell

4. The regulatory regime for adaptive systems — no FDA action, one useful reading [policy]

The FDA published nothing on AI-enabled devices this week; the CDRH AI-enabled device list has not been updated since June 16. The only movement on docket FDA-2026-N-7874 was the SBA's Office of Advocacy soliciting small-business comment on Aug 26 and confirming the Oct 19 deadline. The most useful analysis of the discussion paper came from outside: the observation that a health app giving specific advice and appending "talk to your doctor" probably does not stay in the informational safe harbor — what the output tells the person to do is what determines the category. Separately, ECRI is expanding its adverse-event reporting to include AI errors.

Roadmap implication: the comment deadline is still the highest-leverage cheap action available, and it is now seven weeks out. The ECRI move matters more than it looks: the first durable AI-safety dataset in healthcare is likely to come from incident reporting, not from regulators, and whoever's product is in that dataset first will be explaining it for years. FDA discussion paper · SBA Advocacy

Ripples

1. AI antibody design gets its first blinded report card [wet-lab]

Bradbury and colleagues published the AIntibody benchmark in Nature Biotechnology: 29 organizations, 511 antibodies, three challenges, answers hidden. Best AI affinity maturation 95 pM against 113 pM for conventional phage display — indistinguishable. Best developable de novo design 8.69 pM against 9.2 pM experimental. No single approach dominated; the authors conclude that affinity prediction across tasks remains unsolved and that in-silico confidence does not track experimental performance in any currently predictable way.

So what: this is the CASP moment the field has been asking for, and the score is a draw rather than a rout. A draw is still meaningful — matching three months of wet-lab work computationally has real economics — but it is not the story anyone is selling. Treat any AI-design claim not scored blind, against experimental ground truth, with developability reported, as unscored. Nature Biotechnology

2. The same benchmark, marketed as a rout [wet-lab]

On Aug 26 one participant issued a release headlined "AI Design Surpasses the Best Experimental Result," claiming first, second and fifth of the top five in Challenge 1 at 94.7 pM, roughly 2,000-fold over parental, and the only team placing in the top ten across all three challenges. Every number in it appears accurate. The paper the numbers come from describes the margin over the best experimental clone as statistically indistinguishable.

So what: we are logging this as an item rather than a footnote because it is the single most common failure mode in this sector and it is about to get more common, not less. A true set of facts, a headline the facts do not support, and a peer-reviewed paper sitting right there saying so. The defensive move is procedural: read the benchmark, not the release about the benchmark. Press release

3. Nobody outside Moderna and Merck knows what the algorithm does [Phase 3]

Following last week's Phase 3 win, STAT reported that intismeran's effect may hinge on the software that reads a patient's tumor mutations and picks 34 neoantigens to encode — potentially the first approved medicine whose genetic instructions are chosen by an algorithm. Asked what is in it, cancer-vaccine researcher Alex Rubinsteyn said: "I have no idea. Everyone claims what they have is proprietary intelligence." Some investors think the personalization may not be doing the work at all, and that mRNA itself is priming tumors to respond to Keytruda. Moderna raised $2B in convertible debt on Aug 26, one week after the readout.

So what: an algorithmic component of a therapy that cannot be inspected, replicated, or audited by anyone outside the sponsor is a new category of regulatory object, and effect sizes still have not been disclosed. If this approves, the precedent question — how does a regulator evaluate a black box that selects the active ingredient per patient — matters more than the drug. STAT

4. Generate Biomedicines' AI-designed antibody shows Phase 1 data, by accident [Phase 1]

A conference poster leaked ahead of its Sept 7 ERS embargo, disclosing Phase 1 results for GB-0895 (golukibart), a generatively designed anti-TSLP antibody in COPD: 40 patients, single dose, sustained reductions across four biomarkers, and a half-life of roughly 100 days. The stock hit an all-time high of $20.38 on Aug 25, then fell to $16.57 by Thursday.

So what: this is a genuine clinical-rung datapoint for a computationally designed biologic, which is rarer than the coverage suggests — and it is early, single-dose, biomarker-only, in forty people. A 100-day half-life is the commercially interesting number, not the biomarker panel. Wait for the presented data on Sept 7; the market has already priced and unpriced it twice. Endpoints

5. Capital moves from discovery to execution, and from assets to rent [deal]

Faro raised a $37.3M Series B co-led by Merck's Global Health Innovation Fund and S32, for agentic AI aimed at clinical trial execution, with a stated goal of halving trial timelines within five years; its CEO's pitch is that the startups chasing faster trial documents "just made tiny bits a little bit faster." Separately, Noetik hit the first payment milestone in its five-year GSK deal by getting its OCTO-VC virtual-cell models running inside GSK's own systems — GSK pays subscription-style fees to rent the models rather than buying candidates. Amount undisclosed.

So what: two different bets that the value is downstream of the molecule. Faro prices the clinical timeline, which is where Koller says the irreducible time actually sits; Noetik prices model access as recurring revenue, which is the first evidence that the rent-the-model structure pays out rather than just signs. Both are more informative about 2027 than another discovery-platform round. Endpoints · Noetik

6. Two quiet moves toward measuring clinical AI after deployment [real-world]

ECRI is expanding its adverse-event reporting scope to include AI errors. And the Peterson Health Technology Institute announced on Aug 31 that it will re-evaluate digital diabetes management, arguing the market has changed materially since its 2024 assessment — the one that concluded those tools did not deliver enough clinical benefit to justify their cost.

So what: both are the unglamorous infrastructure that eventually decides which products get paid for. PHTI's 2024 diabetes verdict moved employer purchasing more than any vendor study did; a re-evaluation is a re-opened market, and the vendors that changed their evidence base since 2024 are about to find out whether it counted. PHTI · Fierce Healthcare

Pipeline watch

The scoreboard for whether AI drug discovery is actually working. Two additions this week, both at the clinical end for once.

Company Asset Stage Change this week
Generate Biomedicines GB-0895 (golukibart), anti-TSLP, COPD [Phase 1] Added. Generatively designed antibody; n=40 single dose, four biomarkers reduced, ~100-day half-life, disclosed early via a leaked ERS poster. Full data Sept 7.
Merck / Moderna intismeran autogene (INTerpath-001) [Phase 3] Still zero effect sizes disclosed; ESMO Madrid Oct 23–27 remains the likely venue. Moderna raised $2B in convertible debt Aug 26. New: the neoantigen-selection algorithm is not inspectable by outside researchers.
Insilico Medicine rentosertib (ISM001-055), IPF [Phase 3] No clinical change. Company financials materially de-risk the read: H1 revenue $106.3M, net profit $35.5M, $584.8M cash — this Phase 3 is not funding-constrained.
insitro undisclosed siRNA candidate, MASH preclinical → clinic pending Moved. Koller says the MASH candidate is heading to the clinic by end of 2026 or early 2027 — insitro's first disclosed clinical intent. Liver fat and fibrosis imputed from routine labs and metabolomics rather than biopsy.
Recursion REC-617 (CDK7) and portfolio [Phase 1] No change. TUPELO data for REC-4881 due Nov 2026.
Isomorphic Labs first candidates preclinical No change; nothing published since Jul 16. First-in-human dosing remains the category's marquee test.
Anthropic / Adaptyv Bio minibinder platform [wet-lab], no clinical asset No change. AIntibody is the relevant context: the same rung, now with an external scorer.
Xaira / EvolutionaryScale / Chai platforms preclinical / no clinical asset No change; quiet week across all three.

Zero full FDA approvals of an AI-discovered drug to date. Three FDA approvals landed this week — daraxonrasib for metastatic pancreatic cancer, brepocitinib for dermatomyositis, rusfertide for polycythemia vera — and none of them involved AI in discovery, which is worth saying out loud in a week when the sector's own founder is warning about attribution.


The Signal — Bio/Health, from The Excelsior Group. https://excelsiorgroup.ai/insights/signal/bio/

Don't miss what's next. Subscribe to The Signal - daily AI evolution :
← Newer The Signal — September 2, 2026 Older → The Signal — September 1, 2026
LinkedIn
excelsiorgroup.ai
Powered by Buttondown, the easiest way to start and grow your newsletter.