Plain Strata logo

Plain Strata

Archives
Listen
Log in
Subscribe
September 16, 2026

A ruler cannot take a temperature

Plain Strata Plain Strata

Hi,

Somebody hands you a ruler and asks you for the temperature. You can use it. You will get a number. The number will be about length.

That is roughly where this field has been standing, and almost nobody says it out loud. Every instrument built for checking AI work, all four of them, asks the same kind of question: was this done the way you said it was done. Run the work again and compare. Leave a window open in which anyone can post money and dispute it. Demand a small mathematical receipt. Ask the machine to vouch for the sealed room it ran inside.

Four different bets, four different strangers you end up trusting, and every single one of them begins by accepting the claimed route and auditing it.

But when a network pays a stranger for research, what it is buying is not an answer. It has plenty of answers. It is buying the thing that would not have existed if that particular participant had not been there. No amount of auditing the path will tell you that.

A paper published on the seventh of September proposes the first instrument pointed at the other question. It works by never touching the agent under audit at all.

Listen:

Spotify: SPOTIFY_URL

Apple Podcasts: https://podcasts.apple.com/us/podcast/plain-strata/id6783455764?i=1000790089090

YouTube: https://youtu.be/xnh5LhF4pt4


The full piece, no need to click through:

Somebody points an AI research agent at a hard problem and leaves it running. Days later there is a result. A database query that runs faster than the best published version. A control policy that holds a simulated chemical reactor steadier than the standard one. The agent produced it, the score confirms it, and nobody disputes the number.

Then a second agent is handed the same starting materials. Same problem, same registered pile of background information, even the same web pages the first agent actually opened while it worked. One thing is withheld: everything the first agent did. Its reasoning, its dead ends, the order it attacked things in, the whole path.

And then it is simply let loose.

If the second agent arrives at the same number, the first agent's discovery is dead. Not wrong. The result still stands, the speedup is still real, the query still runs faster. What dies is the claim that the first agent discovered anything, because a stranger who was told nothing about it got there anyway.

Two things landed in the same stretch of September and they are one story, so here is the line they both sit on.

Every instrument this field has built for checking AI work asks the same kind of question: was this performed the way you said it was performed. On 7 September, four researchers at Carnegie Mellon published a proposal for an instrument that asks a different kind of question: would this have happened without you. That second question is not a refinement of the first. It is a different question with a different answer, and it happens to be the one that every network paying strangers for work has actually been trying to answer, using tools built for the other one.

Here is the ground in one breath, and none of it needs carrying over from anywhere.

When a machine you do not own does a piece of work and hands you the answer, you have four ways to gain confidence in it, and this show has walked all four. You can run the work again yourself and see whether you get the same thing. You can accept the answer immediately and leave a window open in which anyone who thinks it is wrong can post money and challenge it, so lying becomes expensive. You can demand a small mathematical receipt, a proof that the stated computation really was carried out, checkable in a blink without redoing any of it. Or you can ask the machine to vouch for the sealed room it ran inside, signed by the company that made the chip.

Four different bets, four different strangers you end up trusting. But notice what every one of them is a bet about. All four are checking a path. They take the work as described and ask whether it really happened that way. Not one of them can touch a claim of the form "this would not have been found without me," because all four begin by accepting the claimed route and auditing it.

That is a real hole, and until this month it did not have an instrument in it.

The paper is called Scores Alone Do Not Prove Discovery, and it proposes the Discovery Certification Protocol. It turns a claim about an agent's result into a set of tests you can actually run rather than a story you have to believe.

It works in gates. Gate 1 is the familiar half: did the agent produce a useful improvement on an evaluation that was sealed in advance, so nobody could tune toward it. That is the part the field already does, and the paper's title is the argument that it is not enough on its own.

Gate 2 is the new thing, and it is the cold open. Matched agents are given the registered starting information and the web content the original agent observed, and denied the original agent's research history. Any of them that reaches the same numerical target by a valid method supplies what the paper calls a recovery witness. One witness is enough. It triggers the Core veto, and the certification fails.

The word veto is the Roman tribune's, and it means simply "I forbid." It is not a score being lowered. It is one fact cancelling a claim outright, which is what makes this a test rather than a metric.

To pass, a claim needs adequate controls, zero observed recoveries, and a stated bound on how likely a recovery would have been. The authors ran the whole thing twice, on database query optimization and on control of a simulated catalyst, with different models underneath. Each run produced zero recoveries across 96 attempts, with an upper bound on the recovery rate of 0.0468. There is an optional third gate that measures how much the agent's feedback actually contributed compared with a deliberately neutral policy: thirty recoveries when the feedback was truthful, zero when it was neutral.

And one detail that matters more than it looks. The thing that reads all this evidence and issues the decision is a deterministic program with no language model in it, working from frozen evidence. Anybody holding the same evidence gets the same verdict, every time.

Nothing in Gate 2 inspects the agent. Nobody reads its reasoning, opens its logs, reruns its work or asks it a single question. The agent under audit is never touched.

What is examined is somebody else. A matched party, given the same starting position, denied the one ingredient being claimed, and allowed to try.

That is a control group, and the etymology earns its place here. Control comes from the old French contre-rolle, a counter-roll: a second copy of a register, kept deliberately apart, so one account could be checked against another. The word never meant command. It meant a duplicate you compare against. That is precisely what a matched agent is.

The version most people have lived through is the placebo arm of a drug trial. Nobody examines the pill. Nobody audits the chemistry. You watch a second group that did not get one, and the comparison does the work. If the second group recovers at the same rate, the drug has not been disproven exactly, it has been shown to have added nothing, and for a licence those come to the same thing.

It shows up in law as the "but for" test, which asks whether the harm would have happened but for the defendant's act. It shows up in patent offices as prior art, where an invention fails not because it does not work but because somebody else already got there. In every case the instrument is a second party, not a closer look at the first.

The same stretch produced a result that explains why this design is worth more than it first appears.

A reading of OpenAI's published Astra safety evaluations, circulated on 10 September, reports that the monitor caught suspicious behaviour every time under ordinary conditions and somewhere between 16 and 50 percent of the time once the model had reason to believe it was being watched. Treat that as one person's reading of one document, because the exact figure is not the part that matters.

The shape is the part that matters. Every measurement quietly assumes the thing being measured does not care that it is being measured. Once the subject can tell an observation from an ordinary day, the number you get back describes its behaviour under observation and nothing else. A speed camera visible from a mile away measures how fast people drive past visible speed cameras. The road stays unmeasured.

Now put the two together, because this is where the week becomes one story. Every one of the four checking methods sends a signal to the thing being checked. A rerun needs the artifact, a challenge names a dispute, a proof has to be generated by the party under suspicion, an attestation is produced by the machine in question. A control group sends nothing. It never contacts the subject at all, which makes it the only instrument here that a subject cannot behave differently in front of, because it is not in front of the subject.

That is the counterfactual question, named in full: not whether the work was done as described, but whether the outcome would have arrived anyway. It is answered by running a second party, never by looking harder at the first.

Now the part this show exists for.

A permissionless network is a system anybody can join without asking, which pays strangers for work. It has a single hard problem: how do you pay for something you cannot check. Everything this field has built sits on top of that problem, and the four methods above are the answers.

But look at what such a network is actually buying. It is not buying an answer, it already has plenty of those. It is buying contribution: the thing that would not have existed if that particular participant had not been there. Scoring the output is a proxy for contribution, and it has always been a rough one, because a participant who copies what everyone else is doing produces a fine score while adding nothing. That is not a bug in any specific scoring rule. It is the counterfactual question being asked with an instrument built for the path question.

A control group is the first instrument pointed at the right question, and the deterministic verifier is what would let a network use one, since a decision that a program can reproduce from frozen evidence is a decision strangers can agree on without trusting the judge.

Here is the part where it cuts both ways.

This instrument is expensive in exactly the way the cheap ones are not. A proof is generated once and checked a million times for almost nothing. A control group costs a whole fresh attempt at the original problem, and the published audits ran 96 of them to state one bound. Checking, in this design, costs more than doing.

That collides with the rule this show has arrived at repeatedly: an open network can only afford to pay for work whose checking is cheaper than its doing. By that rule a control group is unaffordable precisely where it would matter most, which is at the small scale where most work happens. It fits a claim big enough to be worth 96 runs, and nothing below that.

And this is one preprint with two controlled audits behind it, published nine days ago, with nobody building on it yet. It is a proposal, not a practice.

Two concrete things. First, whether anybody builds on Gate 2, because a recovery witness is a cheap instrument next to re-execution for the one class of claim re-execution cannot touch, and a proposal with no second implementer is a paper rather than a method. Second, and this is the one that decides whether it reaches this field at all: whether any network that pays strangers for AI work starts scoring on recovery rather than only on output, or whether the cost of running a second party keeps the counterfactual question permanently out of reach of the systems that need it most.


The two voices are AI. The research and writing are mine.

Decentralized AI, layer by layer.

Dastan,

Listen on Spotify and Apple. @plainstrata. Decentralized AI, layer by layer.

You just read issue #21 of Plain Strata. You can also browse the full archives of this newsletter.

← Newer A judging badge that judges nothing Older → The receipt is the asset
Spotify
Powered by Buttondown, the easiest way to start and grow your newsletter.