Hello,
Picture a laboratory you rent and do not own. The landlord holds the keys to the building, so the chip brings its own sealed booth, and when the work is done the booth slides a signed note under the door saying exactly what ran inside.
That note is what paying customers are buying when they buy trustworthy AI computing. This month it kept its signature and lost its meaning. The booth was still sealed. It was sealing the wrong page.
In the same three weeks a company published the opposite kind of check. No booth, no note. A fingerprint for every one of 80,957 training steps, and a tool that lets you redo any single step on ordinary hardware and see whether your fingerprint matches.
Both are ways of checking that a machine really did the work it claimed. I keep catching myself thinking one of them must be the winner, and neither is. Each one only chooses what to trust instead, and this was the month two of those choices got tested.
So which would you rather believe: a factory you cannot inspect, or arithmetic you cannot afford to run?
Listen:
Apple Podcasts: https://podcasts.apple.com/kg/podcast/plain-strata/id6783455764?i=1000792226223
YouTube: https://youtu.be/5wMlvJbfCJU
The full piece, no need to click through:
Somebody slid a small circuit board between a memory stick and a server's motherboard, cut one wire, and the server stopped noticing it was reading yesterday's data. The board cost about one hundred and fifty-nine dollars in parts. What it defeated was not a password or a lock. It was a machine's ability to tell a distant customer, with a signature, which software it had started up with. That signature is the foundation under one of the two commercial ways to sell trustworthy AI computing.
In the same three weeks, from the opposite direction, a company that has spent years trying to train AI models across machines nobody owns together published a small model carrying something no frontier model ever has. Not the weights, not the data. A public fingerprint for every one of its eighty thousand nine hundred and fifty-seven training steps, and the tool a stranger needs to replay any single step on ordinary hardware and check the fingerprint comes out the same.
Two events, one story. The two live ways to check that a machine really did the AI work it claimed moved in opposite directions in one window. One became something any stranger can re-run on a laptop. The other had its foundation forged on a workbench for the price of decent headphones. That is not coincidence. No method of checking removes trust; each one only chooses what to trust instead, and this was the week two of those choices got tested.
The minimum you need takes one breath.
When you send a question to an AI service, you cannot tell from the answer alone which model produced it. A provider who promised a large expensive model and quietly ran a small cheap one would hand you text that looks fine. That is the gap, and it matters because the whole pitch of an open network of strangers renting out machines is that you need not know or like the operator.
There are few honest answers, and two are shipping to paying customers now.
The first is to re-run the work, or make it possible for somebody else to. Same computation, same result, claim confirmed, nobody believed.
The second is to have the machine testify about itself. The chip carries a key burned in at the factory, runs your program inside a sealed region of memory, and signs a statement saying which program started there. The word is attestation, from the Latin testis, a witness, the same root as testify: the machine bears witness about itself, with a pen belonging to the chip. Picture a sealed booth inside a rented laboratory. You do not own the lab, the landlord holds keys to the building, and the booth is why you can work there anyway.
Both of those moved this month, and they moved apart.
Researchers at KU Leuven, ETH Zurich, Durham University and Google disclosed an attack called DDRop on 14 September. It is not software. It is a board, an interposer, from inter and ponere, to put between, sitting in the slot between a memory module and the processor and keeping up with the memory bus at full speed.
What it does is almost rude in its simplicity. When the processor tries to write something to memory, the board forces an error on the command line, then cuts the single wire the memory module would use to complain about that error. The write is thrown away. Nobody is told. The processor later reads that address, gets whatever was there before, and has no way to know it is holding something stale. The word is exact: stale comes from an old term for something that has stood still too long, used first about beer left sitting in the cask.
This works because of a design decision that is public and deliberate. Confidential computing on modern server chips, Intel's TDX and AMD's SEV-SNP, encrypts the memory a protected machine uses and checks the bytes were not tampered with. It does not check whether the bytes are the newest. That property has a name, freshness, and the vendors gave it up on purpose, because the bookkeeping to guarantee it does not scale to a server holding hundreds of gigabytes. Old encrypted data decrypts perfectly. The seal is intact. It is sealing the wrong page.
Think of a vault with a tamper-proof logbook. Every entry is genuinely signed and none is forged, so every check passes. Somebody removed today's page, and yesterday's page underneath is still perfectly sealed. The vault reads it and believes it is current.
With that one move the researchers dropped the writes that set up a protected machine's page tables, left their own data there instead, mapped their own machine onto a victim machine's memory addresses, read it, flipped it into debug mode to see the data in the clear, and restored it with no trace. Then came the part that matters most here. They overwrote the launch measurement, the fingerprint a machine signs to prove to a remote customer that it booted in a known state. Forge that and an attacker's machine presents itself as a trusted one.
Intel and AMD both answered that physical attacks on server memory sit outside their published threat model, and Intel does not plan to assign this a tracking number. The answer is not dishonest and it is not sufficient, because customers buy confidential AI inference precisely so they need not trust whoever is standing next to the machine. That is why this is a story about what a check decides to trust. This branch lands its trust on a fabrication plant, and somebody just showed the last few centimetres of wire are part of that bet.
Now the other direction. On 15 September, Gensyn published open-1b, a model of about 1.61 billion parameters trained on 400 billion words of permissively licensed text. It is small and its scores are modest: 25.4 on a standard evaluation suite against 31.9 for a comparable open model. Nobody should care about the model. The interesting object shipped beside it.
Every one of the 80,957 training steps has a published canonical hash covering the batch of data, the parameters, the optimizer state and the gradients. A hash is a fingerprint: run any block of numbers through a fixed grinder and out comes a short value that changes completely if one bit of the input changes. So an audit becomes mechanical. Download the harness, pick a step, load the checkpoint from just before it, replay the step, hash the result, compare.
That is a much bigger claim than it sounds, because of the oldest obstacle in this field. Two honest machines running the same computation do not produce the same numbers. Floating-point addition is not associative, meaning a long list of numbers added in a different order gives a slightly different total, and thousands of processor cores race each other and finish in a different order every run. So "re-run it and compare" has never been available: any difference could be cheating or could be physics, and nobody can tell which.
What Gensyn built is not the hashing. It is the machinery that makes the replay possible at all: a library of reproducible operations fixing one summation order for every sum, one multiply-and-add convention, and identical handling of the smallest numbers, so matrix multiplication and gradient reduction produce the same bits on a consumer graphics card, on an x86 or ARM processor, and on Apple Silicon. The stream of training data depends only on the starting seed and the corpus, not on how many machines were used, so an auditor with one machine can reproduce a step that ran across many.
The price is stated plainly, which is rare and worth crediting. About five percent model FLOPs utilization, roughly five times slower than an optimized standard stack on the same hardware, 27.8 days of active training on 48 top-end accelerators. Determinism is not free. It is bought with throughput, and the bill came to a factor of five.
The design move underneath it is the part worth carrying away. They did not decentralize the training, which is what this company is known for trying. They decentralized the verification. Replaying a whole run is still out of reach for one person, so the scheme is collective: many independent auditors each certify individual steps and together cover the run. The ledger of who verified what carries names, not payments. No token, no reward, nothing staked. The offer is your name on the record of the first fully audited training run.
The honest limit is published in the same post: this tool cannot audit Llama or GPT. Reproducibility has to be designed into a run from step one, never bolted on afterwards, which is an older rule in a new room. A thing has to be made countable before it can be made accountable.
One small item in the same week says more than its size suggests. On 23 September the same company announced two research hires, both from the study of proper scoring rules and information elicitation: the mathematics of paying people so that honest reporting is their best move when you cannot check what they reported.
Read that against what they had just shipped. Bit-exact re-execution is the strongest check available, they built it, and then staffed up on getting truth out of parties whose work you cannot re-run at all. That is a company telling you, through its hiring, where the boundary of its own best instrument sits. Re-execution works when the work is deterministic and somebody pays the five-times tax. Outside that you are back to pricing a claim rather than confirming it. Which is why these announcements belong in one episode: the strongest check in the field arrived with a map of its own edges.
Name the shape once, in full, because it is the most useful thing here and it explains why both events landed in the same window.
No verification scheme gets rid of trust. Each one moves it somewhere you like better, and judging one is mostly a matter of naming where it lands. Re-execution puts your trust in determinism, in arithmetic behaving identically twice. Hardware attestation puts it in a fabrication plant. Cryptographic proof puts it in mathematics, and sometimes in a setup ceremony held years ago by strangers. Stake-weighted scoring puts it in whoever holds the most stake.
You can run this test in thirty seconds on anything new. Skip the mechanism, find the relocation, ask whether you prefer that party to the one you started with.
This month ran the test on two bets at once, from opposite ends. Somebody attacked attestation's landing place with a board and a cut wire, and the vendors replied, correctly, that the landing place was never advertised as covering it. Somebody else manufactured re-execution's landing place from scratch, engineering determinism into existence where the hardware does not supply it, and published the invoice. Neither is a verdict on which method wins. Both say the same sentence twice: a check is exactly as good as the thing it decided to believe.
The pattern is older than computing. An audit relocates trust from a company to an auditor, a notary from the person signing to the state. Audit comes from the Latin audire, to hear, because an audit was once a hearing, the accounts read aloud to someone who could not read them. The interesting modern version is the opposite of a reading: an audit you perform by doing the work again yourself.
The honest implication is uncomfortable for the side winning commercially. The cheap-to-sell check, a signed statement from a sealed machine, is what paying customers have been buying, because it also hides their data from the machine's owner. It now has a hundred-and-fifty-nine-dollar counterexample and no tracking number. The expensive check, replaying the arithmetic, is the one that got stronger, in a form with no business model attached: no token, no payment, names on a list.
Do not oversell either. A physical attack needs hands on a machine, a real barrier for a large cloud and much less of one for a rack in a building nobody has audited. And a collective audit with one auditor is a claim again.
Which points at the open question, and it is small and answerable. Will anybody outside the company that built it actually run the audit? If the ledger fills with outside names, decentralized verification has its first working instance, on a model small enough not to matter and a method large enough to. If it stays empty, this was a very good paper.
The two voices are AI. The research and writing are mine.
Decentralized AI, layer by layer.
Dastan,
You just read issue #24 of Plain Strata. You can also browse the full archives of this newsletter.