Hello,
Picture a sealed booth dropped into the middle of a landlord's room. The wall is opaque, the lock is perfect, and the landlord holds every key to the building and cannot use any of them on this door.
He still owns the floor, and the booth stands on it. Every time somebody inside walks to a filing cabinet and opens a drawer, the boards creak. He cannot read a page. He learns which drawer opened, in what order, and how many times.
A research paper this year showed that for one particular program, that is enough. The program turns your sentence into numbers before an AI model ever sees it, and the order of its drawers is the order of your words. Nothing was decrypted. The prompt came back anyway.
I keep catching myself thinking privacy is a question about what you said. It is also a question about how a machine moved while it listened.
So which would you rather protect: the letter, or the footsteps of the person carrying it?
Listen:
Apple Podcasts: https://podcasts.apple.com/kg/podcast/plain-strata/id6783455764?i=1000792611924
YouTube: https://youtu.be/Hx59Qe3LEhg
The full piece, no need to click through:
A researcher rents a server, the kind a cloud company sells for private AI work. The selling point is that the company cannot see inside. The memory is encrypted with a key that is generated inside the processor package and never leaves it. The operator who owns the building, the rack, the power and the cooling can stand in that room all day and see nothing but scrambled bytes.
Then the researcher, playing the operator, reads the user's prompt back. Not a guess at the topic. The words, in order, more than ninety percent of them, from a single run. Credit card numbers came out whole. So did social security numbers. No byte of encrypted memory was ever decrypted, no key was ever stolen, and the cryptography did exactly what it promised.
That result is called TDXRay, and it appeared at the IEEE Symposium on Security and Privacy in 2026, after disclosure to Intel and to the large labs in November 2025. It is worth thirty minutes because of what it teaches, which has nothing much to do with Intel and everything to do with a distinction most people have never had a reason to notice.
Encryption hides what a value is. It has never hidden which address was touched.
That is the whole episode. A sealed machine still has to reach into its own memory, and reaching is a physical act on hardware it shares with its landlord. The sequence of places it reaches for is not a secret the encryption covers, because the encryption was never applied to the sequence. It was applied to the contents. And for one particular program, the one that turns your sentence into numbers before an AI model ever sees it, the sequence of places is enough to rebuild the sentence.
Three things have to be on the table.
A virtual machine is a whole computer built out of software, running on somebody else's physical computer. The program that creates and supervises it is the hypervisor, and it is the landlord's program. Normally a hypervisor can read every byte of every machine it hosts, the way a landlord with a master key can open any door.
A confidential virtual machine is the answer to that. Intel's version is called Trust Domain Extensions, or TDX. The hardware encrypts that machine's memory with a key held inside the processor, and refuses the hypervisor's reads. This is the enclave idea, from the French enclaver, to lock in, from clavis, a key, and also the map shape: one country's land entirely surrounded by another's, which is exactly what a region of a machine the building's owner has no jurisdiction over amounts to.
And attestation, from testis, a witness, the same root as testify: the machine bears witness about itself, with a pen that belongs to the chip. It signs a statement saying which software launched inside the sealed region, so a customer far away can check the room before speaking into it.
Put together, the picture is the one worth holding: you do not own a laboratory, so you rent space from a landlord who holds every key, and you drop a sealed booth into the middle of his room. One hatch, and a window that is a mirror from his side. You fired the landlord and hired a factory.
This episode is about the mirror. It turns out you can learn a great deal by watching a mirror shake.
Side channel. Channel is from canalis, a pipe or a groove. A side channel is the groove running beside the pipe, and the water in it was never water you meant to send. The signal is real, it carries information, and nobody designed it, which is precisely why nobody guarded it.
Cache. From the French cacher, to hide. A cache is a hiding place: a small, fast pocket of memory close to the processor, holding the handful of things it touched most recently so it does not have to walk all the way out to main memory again. The name is now the joke. The hiding place is what gives the secret away.
Probe, from probare, to test, to try, to find good. The same root as prove. In this story the probing is done by the attacker, and it is the identical act: try it and see what comes back.
Oblivious, which matters because it names the defence. From obliviscor, to forget, built on levis, smooth: to smooth over, to wipe flat, the way you wipe a wax tablet clean. A data-oblivious program is one whose surface has been smoothed so the input leaves no mark on it at all.
Token, from the old word tacen, a sign or a mark. A token is a sign standing in for a thing, and a tokenizer is the program that swaps your words for signs.
Forget software for a moment and look at the silicon.
Main memory is a set of chips on sticks, some distance away from the processor across a bus. Fetching from there is slow in processor terms, so the chip keeps caches: small arrays of very fast memory on the same piece of silicon as the cores. Data moves in and out of them in fixed blocks of sixty-four bytes, called cache lines. Memory is also divided, for the operating system's purposes, into pages, usually four kilobytes each.
Now the thing that makes this episode possible. The cache is one piece of hardware, and both machines on that chip use it. So does the memory controller, the toll gate sitting between the cores and the memory sticks, through which every read and write passes. TDX's whole trick lives at that toll gate: pages marked as belonging to a protected domain, encrypted on the way out with a key made inside the chip, refused to anyone outside the domain.
Refused. Not hidden. The landlord's program asks for a line, and the hardware says no. But a no is an answer, and the timing of the no, and what happens to the shared cache around it, are observable events on hardware the landlord also owns.
Here is the plainest version. The sealed booth has an opaque wall, and the wall is genuinely opaque. But the booth stands on the same floor you are standing on, and when somebody inside walks to a filing cabinet, the floor creaks. You cannot read a single document. You can learn which drawer was opened, in what order, and how many times.
The story has four steps, and each one causes the next.
Step one. Intel said so. Intel's published threat model for TDX explicitly excludes microarchitectural side channels. That is not a slip. Blocking them is genuinely hard and costs performance, and the technology was built to stop a hypervisor reading memory, which it does. So the gap was documented from the start, in the way a limitation is documented when nobody has yet shown what it costs.
Step two. Somebody measured the gap. TDXRay is a Linux kernel module running on the host, inside legitimate host interfaces, needing no cooperation from the machine it watches and no physical access to the building at all. It combines four signals, and the reason there are four is that each one alone is too coarse or too noisy. SEPTrace uses a host interface TDX provides for blocking guest memory translations: block every page in a region, resume the machine, and wait. Each fault names exactly which page was just wanted. That gives a page-level trace, four kilobytes at a time, deterministic, no timing involved. Load+Probe goes finer: reading a second address that points at the same physical line takes measurably longer if the sealed machine has that line cached, so a stopwatch reveals not just that a line was touched but whether it was read or written. TSX-Probe does the same thing with no stopwatch at all, by wrapping the read in a hardware transaction that aborts on a conflicting cache line, which turns a noisy measurement into a clean yes or no. And MWAIT-Probe is the metronome: an instruction that pauses until a specific physical address is accessed, which lets the watcher line its measurements up with the sealed machine's execution rather than sampling blind.
Stack those and you get what the paper delivers: a cache-line-granular trace of memory accesses inside an unmodified confidential machine. Sixty-four bytes of resolution, from outside, through encryption that was never broken.
Step three. Point it at the one program that betrays everything. This is the level worth going all the way down to, because it is the whole point of the episode.
Before a language model does anything with your sentence, a tokenizer converts it into numbers. Models do not read letters; they read identifiers from a fixed vocabulary, often around 128,000 entries. So the tokenizer holds a hash map: a lookup table where a piece of text is run through a fixed public function that produces a number, and that number says which bucket of the table to look in. Several pieces of text can land in the same bucket, so each bucket holds a small chain of entries, and the lookup walks along that chain comparing until it finds the match.
Now hold two facts together. The table lives at addresses in memory, and each entry sits somewhere inside some sixty-four-byte line. And the hash function is public, because it ships with the model. So an attacker can compute, in advance and offline, exactly which lines in that table a given word's lookup would touch, and in what order.
The attack is then mechanical. Localize the table in the trace. Record, cache line by cache line, the walk the tokenizer performs. Replay the deterministic lookup procedure against your own copy of the same table, matching the observed walk to the only word that produces it. Stitch the recovered tokens back together in order, and you have the prompt.
Nothing was decrypted. What leaked was the shape of a walk across a table whose layout was never secret. Reported similarity is over ninety percent across both Llama 3.2 and Gemma 3, from a single trace, on three generations of Intel server processors.
And now surface, because this is the sentence the dip was for: the sealed booth kept every promise it made about contents, and the contents were never where the secret lived.
Step four. The part that should worry the field. Tokenization runs on the processor, before the model runs at all. Every dollar spent on confidentiality for the graphics card, the encrypted link to it, the card's own signed statement, the whole second sealed room this show described when it walked through hardware attestation, contributes exactly nothing here. The prompt is gone before the expensive machinery is reached. A chain of protections is only as strong as its first link, and the industry spent two years hardening the third.
Name it plainly: the pattern is the payload. When a system protects the contents of its messages and leaves the shape of its activity in the open, the shape is often the more useful of the two, because the shape is what a system does rather than what it says.
Military traffic analysis is the oldest clean instance. You do not need to break a code to learn that one headquarters suddenly sent forty times its usual volume of traffic to one coastal unit at three in the morning. The messages stayed secret and the attack did not.
The distinction survives in law, which is the tell that it is real. A wiretap needs one kind of warrant and a record of who called whom needs a weaker one, on the theory that the second is only metadata. Anyone who has looked at a month of somebody's call records knows how thin that theory is.
Domestic electricity meters carry the same lesson at household scale: read the draw finely enough and the appliances identify themselves, so the utility learns when you slept and when you showered without reading anything but a number going up and down.
And a library's borrowing log tells you what a person is thinking about while the books stay closed.
In every case the protected thing was the content and the leak was the activity, and in every case the defence is the same in kind: make the activity uniform, so it stops carrying information. That is what data-oblivious means, and why the word is the right one. Not hidden. Smoothed, until there is nothing to read.
This show has a rule it applies to every verification and privacy scheme it meets: no scheme removes trust, each one only moves it, and the analysis is naming where it lands. Hardware attestation relocates trust onto a fabrication plant. What TDXRay demonstrates is that the confidentiality half of that bet relocates trust onto something softer than a factory: onto a threat model, which is a document, written by the vendor, describing which attacks it intends to stop.
Read that honestly and it cuts both ways. Intel is not being evasive when it says microarchitectural side channels are out of scope, because that was written down in advance and the cryptography did hold. But a customer buying confidential inference is buying it for exactly one reason, which is that they do not want to trust whoever stands next to the machine. When the answer to a break is that this particular way of not trusting the operator was never covered, the product and the reason for buying it have quietly come apart.
There is a harder thought underneath. Every privacy technology so far protects what you said. None of them protects that you were thinking in a particular shape. The tokenizer leak is uncomfortable in a specific way: what leaked was not your data sitting in storage, it was your sentence in the act of being understood, caught in the instant between language and number. A machine that has to look something up has to look somewhere, and looking somewhere is a movement in the world.
Walk one real request through, with every piece named as it arrives.
You send a prompt to a private inference service. The service runs in a confidential virtual machine on a rented server. Before it will talk to you, it signs an attestation: this exact software launched inside the protected region, this is genuine silicon from this manufacturer, and your random number is inside the report so you know the statement is from today, not from last year. You check it against the fingerprint you computed yourself in advance. The room is established once, and everything after rides on that one check.
Your prompt arrives encrypted, and is decrypted inside the sealed region. Every page of it passes the memory controller's toll gate, encrypted on the way out to the memory sticks with a key the operator cannot reach. The hypervisor asks for those pages and is refused.
Then the tokenizer runs. It walks its hash table, bucket by bucket, cache line by cache line, once per piece of your sentence. The walk happens on shared silicon. The operator's kernel module blocks pages, catches the faults, probes the cache with a transaction that aborts on conflict, and aligns its measurements with an instruction that waits on a physical address. It writes down the walk.
The tokens then go to the graphics card, whose memory is also a protected region, over a link that is also encrypted, with debug interfaces and performance counters switched off. All of that works perfectly. The model answers. The answer goes back to you encrypted.
Meanwhile the operator replays the walk against a public table and prints your prompt.
Every protection in that paragraph did its job. The prompt is on the operator's screen.
Three things, all answerable.
Does any vendor selling private inference ship a data-oblivious tokenizer? The researchers built one for a full 128,000-token vocabulary and report that the overhead is acceptable, since inference time dominates anyway. That mitigation needs no help from Intel and no new hardware, which makes it the cleanest possible test of whether these products are engineered to a threat model or marketed to one.
Does the hardware fix arrive? Two of the four primitives work because the cache does not distinguish lines with the same physical address under different encryption keys. Putting the key identifier into the cache tag would let them coexist and kill the signal. That is a silicon change, which means years, and it will have a cost somebody has to be willing to pay.
And the one this show keeps returning to: is detection a defence? A monitor inside the sealed machine can watch its own page faults and cache misses and notice that it is being traced, because the tracing leaves a footprint. But noticing you are being watched by the party who owns your power cable is a strange kind of protection, and it turns a security property into a race.
The two voices are AI. The research and writing are mine.
Decentralized AI, layer by layer.
Dastan,
You just read issue #25 of Plain Strata. You can also browse the full archives of this newsletter.