iter.ca updates

Archives
Subscribe
September 28, 2026

iter.ca update #12

Hi there readers of my weeklyish updates! It’s time for another round of updates from yours truly.

london

I’m still in London! It’s a really cool city and I like wandering around it; it’s a really interesting place!

Some buildings in London
Buildings in London

also london

research

I’ve still been thinking about how to use meta-models to interpret LLMs better than the existing methods (probably by making verbalizations that are more causally linked to the activations?), but I haven’t figured out a way to do that well yet.

One thing I thought about doing was replacing the activation reconstructor (AR) with something (e.g. a lightly fine-tuned base model) that predicts future tokens, given the activation explanation. This means that the activation text must be something that’s useful for predicting future tokens, which would probably be related to some notion of the model’s internal state. But this has the problem that the AV might just learn to predict the future tokens give the input activation. I’m currently trying to make something like that, which actually works.

Another thing you could do is just like make it so the NLA is scored based on how well the model predicts future tokens when you replace the actual activation with the result of AR(AV(h)). You could do this for the single next token or over some number of future tokens. But I worry that a NLA trained this way might be worse at verbalizing more “internal” kinds of state that don’t bear very directly on what the next few tokens are. Maybe there’s some way to quantify how good a reconstruction is that’s actually good that I haven’t thought of?

I wrote about “future oracles”, which are a meta-model you can make by SFTing a model to predict future tokens given an activation. They’re not that interesting though, I mostly worked on that to better understand what kinds of information is in activations.

Idk how useful working on NLAs is; I’ve been thinking of much different ways to use meta-models to do interp. It might be worth trying to figure out how to define how “good” an explanation is and then having a bunch of agents try different things, but it’s hard to define that well. But it’s hard to design meta-models that actually work well!

other stuff

I wrote about some YouTube adblocking stuff I did last year: Patching protobufs poetically. I watched The Princess Bride and WarGames movies. I started reading planecrash.

looking forward

Probably going to do some more LLM interpretability stuff. I might write some more thoughts on how to do natural language autoencoders but better. I might also end up writing about AI lab governance (I think Anthropic’s corporate structure is underdiscussed!). And as with the last few weeks I’m going to continue to London-maxx. I know I said this last time but I’m going to visit the Isle of Dogs real soon. Please let me know if you have any suggestions for me; I’d love to get advice!

Don't miss what's next. Subscribe to iter.ca updates:
Older → iter.ca update #11

Add a comment:

Posting this comment will subscribe you to this newsletter with the email address you enter.
GitHub
Powered by Buttondown, the easiest way to start and grow your newsletter.