Tracing the rogue ideology of the frontier labs to their product choices

This is a really long newsletter, for which I apologize, but hopefully there's enough detail in there for people who have not been following this stuff closely to get up to speed. If you know, for instance, what a 'transformer' is, some parts are skippable.
The question of "AI Safety" is one that has taken on new prominence in media and government. This is due to a series of admissions—or, depending on your level of cynicism, PR stunts—from the so-called AI "frontier labs" OpenAI and Anthropic, admissions that their best models, when set to work on internal cybersecurity testing, had contrived to hack external sites in various ways on multiple separate occasions. Depending on who you ask this is either a sign of unimpressively slapdash and incompetent internal network security and testing practices or a sign that the models that these labs are producing are at imminent risk of loosing their creators' grip and becoming malevolent forces free in the world. The frontier labs broadly take the latter view, and have advocated for a coordinated slowdown in AI development that would involve otherwise inappropriate collusion between the big players in the space to agree to stop work on ever-better models until some not yet fully specified set of future conditions obtains. This is, as has been broadly noted, strange on the face of it—what other industry do you know that has voluntarily offered to pause its march towards profitability and lucrative IPOs—and if you haven't been more-or-less marinating in the culture that birthed these labs it might seem frankly inexplicable. If you have been, what has been happening comes into focus, and it is neither the story of incompetence nor the story of imminent robot apocalypse but instead a narrative driven by a highly insular culture with some deeply unusual embedded assumptions taking an approach that naturally follows from those assumptions and seeing its biases confirmed as a result of it.
What is happening within the labs is that they (or at least, many of their earliest and most important technical contributors) have a strong prior belief—at the level of ideology, with a certain flavor of religious commitment—that AI is inevitably going to happen, will inevitably exceed human capabilities, and is inevitably going to have its own more-or-less alien internal motivations. They further believe that these motivations, if not carefully sculpted to match human intuition, will lead to the AI "going rogue", and that without their explicit intervention self-improving rogue AI will be the downfall of humanity. This guided their early work and now, after the early approach that matched these intuitions did not pan out, they are working from the premise that modern AI—LLMs—will match the intuition soon enough, and in practice, without even necessarily meaning to do so, are working to make their assumptions true.
To understand these assumptions, and how they led to the decisions and approach of these companies, it is necessary at some level to understand Bay Area AI culture, and the role of a man named Eliezer Yudkowsky. This is a massive topic, by turns ridiculous, infuriating, horrifying, and prurient. It is challenging to get a general handle on, not least because the people involved are overwhelmingly prolix, and I couldn't begin to do it justice in a single essay. If you want to dig in deep there are books, excellent recent articles and blog posts aplenty. We don't need to get terribly deeply into it to map how it affects the culture and approach of OpenAI, Anthropic and, at least to some degree, their American competitors like xAI and Google's DeepMind.
The very short form (all of this got hashed out over the course of years in tens of thousands of posts on the forum LessWrong as well as in various epic-length texts including the nearly three-quarters-of-a-million-word Harry Potter fan-fiction) is that there is an overlapping set of people who describe themselves variously as "rationalists" and "effective altruists" who have convinced themselves with what they insist is not religious certainty that we are trundling inexorably down the road towards AI that is not just smarter than humans, but so much smarter than humans that we will be unable to understand or control it. This is not a new idea—per wikipedia it dates to 1965, but the foundation texts of its modern proponents started to appear in the early 1990s—but it is one that got codified and reified by the rationalist/EA communities in the first decade and a half or so of this century. Particularly, both self-described EAs and self-described rationalists (groups which, I should note, insist that there are strong and obvious differences between them, and that they are often at odds) believed that by using the techniques of Bayesian statistics you could establish by process of rational induction at least reasonably strong statistical bounds on both the eventual likelihood and the consequences of the creation of (in their words) "superintelligent" AI.
The statistical commitments that are required to convince yourself that this is a matter of factual inference arrived at by rigorous induction tend to give actual working statisticians fits, but they were sufficiently convincing to many in and around the tech world that Eliezer Yudkowsky turned his concerns into a nonprofit institute named MIRI that raised substantial sums from Silicon Valley techies (both executives and the rank-and-file) looking for a philosophically congenial home for their charitable dollars. Yudkowsky’s work inspired both acolytes and opponents, with factions like "doomers" and "accelerationists" emerging to take different positions on the desirability of the inevitable superintelligence. With intellectual heft provided to the broader movement and its debates by fringe British academics named Nick, Yudkowsky and MIRI's campaign and its offshoots gained sufficient profile to engage some of Silicon Valley's wealthiest and most connected insiders, like Sam Altman, Peter Thiel, and Elon Musk.
The premise that engaged those wealthy VIPs had been shaped, by that time, into something very specific and distinctly strange. It held that superintelligent AI was essentially inevitable. It held that the very most worthwhile thing a person could do, from a utilitarian perspective, was to act in the interest of the quadrillions of future people who would be most affected by that superintelligent AI. It held that if the process of development of that superintelligent AI was not handled in an incredibly precise and careful way, the AIs we ended up with would nigh-inevitably extinguish humanity. Therefore, the very most important thing that we can do right now is try to develop NON-evil superintelligent AI as quickly as possible. I should note, by the way, that Eliezer Yudkowsky has subsequently renounced that final leap in logic, that the right answer is to build AI now so it can be controlled, and is now more prone to suggest things like nuking data centers as part of an at-all-costs effort to stop AI development, but the intellectual lineage on both sides runs, broadly, through him, and the groups funding construction of datacenters remain oddly congenial with the groups hoping to immolate them.
Others have done a better job than I will of laying out the details of this worldview, should my explanation feel insufficient. But I will do my best, briefly, to explain it, and I think the best way to do it is to talk about the Paperclip Maximizer. The Paperclip Maximizer is a thought experiment devised by Oxford-based philosopher Nick Bostrom. It asks us to imagine a future Artificial Superintelligence—that is, an intelligent agent which is able to best human performance at any possible task requiring intellectual effort. This ASI is asked by the humans which created it (or created it in part) to produce as many paperclips as possible. Because this ASI is fundamentally a computer program, it will perform the task that it is given without question. In this case, it will take the exhortation to produce as many paperclips as possible entirely literally, exhausting the galaxy's resources, cajoling its human minders into giving it ever more power, and eventually, in its single-minded pursuit of well-clipped office memos, destroying all human life in the process. It will not do this because it is evil, or angry at humans, but because we are irrelevant to its goal, and its alien intelligence is inhumanly single-minded and blankly sociopathic in its goal-directed behavior.
These ideas were and are phenomenally influential. They directly inspired the creation of OpenAI, as a nonprofit, with the founding goal of creating artificial superintelligence aligned to human values and uncorrupted by the pressures of capital (that got complicated). They even more directly inspired the schism between OpenAI and the people who went on to found Anthropic. The set of ideas which revolve around rationalism, effective altruism, and the analogy of the paperclip maximizer are central to the self-conceptions of those companies and many of the people who work there, most particularly the people who work in AI Safety and "alignment" (a term from that world that we'll get into further below) and the PhD researchers working on model training.
The central goal of Anthropic and (to a lesser degree) OpenAI is to create a superintelligent AI which is aligned to human values. That is, they believe that we are on the cusp of creating machines that vastly surpass us in every field of intellectual endeavor, are (or at least, “are” with enough plausibility that they tried to convince the pope of that) best regarded as conscious intelligences, and which must be created with an inherent value system which leads them to act, of their own volition, in ways that will be beneficial for the humans of the far future. Everything that seems puzzling or incoherent from the outside about their product decisions, testing plans, or public statements falls out of these tenets.
When OpenAI was founded in 2015, there was no such thing as a large language model. The algorithm which underlies large language models, the "transformer" algorithm, was not invented until 2017. The teams at OpenAI—including notably the one led by Dario Amodei, now CEO of Anthropic—were working on a different technology. They were focused on models which used a technique called reinforcement learning. In machine learning, "classical" reinforcement learning (which is more-or-less the kind that they were interested in) involves a system which represents a state of the world and a set of possible actions to take in that world. The state is represented by a long string of numbers known as a "state vector" and the set of possible actions is represented by a different string of numbers known as a "policy vector". Approximately, the model learns to associate different states of the world with the action in the policy vector which will lead, eventually, to the model maximizing its reward.
This is an approach that maps closely to how the brain works—one of the most impressive papers in the history of computational neuroscience linked a version of reinforcement learning called TD-learning which was capable of learning to play backgammon to the firing rates of dopamine neurons in the brain—and seemed to many people in 2015, including Google's DeepMind, to be the most promising approach to making progress towards 'AI'. Just as important, much of the thinking that the rationalists and their fellow travelers had done about the dangers of 'AI' had been influenced by the way that reinforcement learning works and fails. The policy vector in an RL system can be thought of as that system's internal motivation; it is the things that it wants to do in order to achieve its goals. The paperclip maximizer example can be thought of as a RL system where the reward it is seeking—which is linked to the number of paperclips it can produce—is a direct reflection of the expressed desires of its human creator but the specifics of the policy for how it gets there are not, because the reward is underspecified. The intervening steps it takes as a product of its internal motivation lead it to do things which are unwelcome at best to its human creators.
So OpenAI—specifically Amodei's team—set to work trying to remediate these problems. They created an approach called "reinforcement learning from human feedback" where at a given step the RL agent would present two different plausible actions from its policy vector and a human reviewer would select which one was most aligned with what the human believed the appropriate behavior to be. Eventually, the agent would learn a policy that maximized its reward without ever doing something that a human reviewer would regard as inappropriate or reward hacking. The alignment problem, with the paperclip maximizer as its asymptote, seemed solved.
There was another problem, though, which was the kind of algorithms OpenAI was working with (called deep reinforcement learning) weren't very good at solving real-world problems, and were improving only slowly. The problem with reinforcement learning as an approach is the larger the system's state vector is—that is, the bigger and more complicated the world it lives in—the larger its policy vector needs to be and the more data and time is required for the model to "converge", or settle down to a stable reward-maximizing policy. These data and scale problems are so acute that in practice deep RL systems have only ever reached convergence in environments where the state and policy vectors are highly constrained. They are excellent at playing relatively simple video games, where the number of actions and outcomes are inherently limited, and good (if you throw truly enormous amounts of computing power at them) at some useful tasks like protein folding where, while the problem is complicated, the state and policy at any given point can be constrained to be relatively small. But nobody has had much luck getting them to operate in unstructured real environments.
In the meantime, across town at Google, researchers made sudden and surprising progress with a completely different approach to machine learning. The transformer algorithm is a kind of artificial neural network—there are "nodes", or neurons, which have weighted connections to each other, and those weights are learned by training it on the statistical correlations in real data—but it is in some sense the hot rod of neural networks. Everything unnecessary to the task of representing statistical correlations (much of the internal structure, filter mechanisms, hierarchies) has been stripped away, leaving a network that seemed at first glance to many too simple to possibly work. But in its simplicity it turned out to have two properties that would upend the world of machine learning. First, because it had no hierarchical internal structure, it was capable of learning statistical relationships between parts of a training sample (portions of an image, say, or words in a text) that were arbitrarily far away from each other. Where a traditional deep network operated within limited windows—you can see a visual representation of this in Google's early Deep Dream generative models, where it would tile an image with dog heads because it had no sense of overall image structure—transformers could operate across the whole sample. In particular, when working with a string of text either as input or as output, they could represent connections between the first word (or fifty words) in the text and the fifty-thousandth. Second, because of the transformer's architecture, models could be trained on arbitrarily large amounts of data and just keep getting better, where earlier approaches would "asymptote", with improvements plateauing after a relatively small (in ML terms, so small could mean "hundreds of thousands") number of training samples.
The transformer algorithm, then, allowed for language models—models which predict the next word in a text—that were hugely grander in scale than any that had been attempted before. Models that were trained on a substantial portion of all the written words ever produced (a data set that Google helpfully already had on hand). This is the LLM. What nobody precisely expected was exactly how good LLMs would be. Not only were they almost immediately capable of feats of translation and summarization that had seemed out of reach even a year earlier, but they were basically suddenly able to produce effortlessly fluent human language.
The results that Google achieved with transformers were immediately impressive enough that pretty much everybody in the industry started looking into them—the autonomous car world, which is where I was working at the time, started seeing transformer-based approaches within a couple of years even though the compute needs of transformers are incredibly difficult to accommodate in a car; in subfields with the ability to run in data centers adoption was faster—and OpenAI was no exception. They didn't abandon their previous work right away, but the efficacy of transformers and LLMs was too much to ignore.
Google's LaMDA LLM got rapidly ever-more-effective at producing fluent language (but, importantly, ONLY at producing language—it could not yet search data or use tools); OpenAI played catch-up with GPT; Google developed a measured release strategy involving weighing the dangers; Sam Altman seized the opportunity to YOLO and released an underbaked and barely constrained ChatGPT on the world. All hell broke loose. Around this time (before the release of ChatGPT, but possibly because of early discussions about making models publicly available) Dario Amodei, unhappy by most accounts with OpenAI's insufficient attention to safety and alignment, left to start Anthropic.
What I think is most important to understand about that time is that the LLM released by OpenAI had essentially nothing to do with what they were building and what they had believed that AI would look like. Their bet, on both the promise and the danger of AI, had been wrong. Reinforcement learning, as they conceived it, could not compete.
One of the key differences between the models that they'd been building and LLMs is that LLMs lacked a policy vector. When you train a machine learning model, the training process is guided by what's called an "objective function" or "loss function". This is a function which compares the desired output to the real output, and gives the difference between the two; it is called a "loss" function (usually) because the ideal outcome of training is for the actual output to have lost zero of the information that characterizes the desired output. So you have a number, loss, that you are trying to minimize. In the reinforcement learning case, that number is the difference between the reward achieved at the end of a simulation and the maximum possible reward that could have been achieved; what the algorithm does is partial out credit for achieving that reward to the policy steps taken at every point, and update the policies in the direction of each individual policy element contributing a little more to the reward. This iterative process is how you end up with a policy which has converged.
In the case of a pure LLM, the loss function is evaluated against a piece of text for each token (a "token" in a LLM is not quite a word but you can usually think of it similarly; that definition will do for us here) produced. The model ranks the next tokens it could produce in terms of how likely it believes they are to be the "correct" next token, and the loss function is how much that probability deviates from certainty for the actual next token in the training data. There are other ways to look at this—one useful one is to talk about it as the model's "perplexity", or how many different tokens it finds more-or-less equally likely to be candidates for the "true" next token—but the important thing to understand is that it is purely a measure of how well the model predicts the next word given all the words that came before (or, at least, "all" up to the amount it is able to represent at once given its architecture). There is no longer-term goal and no particular internal representation of the state of the world. You can sort of talk about the previous tokens in the model's window of visibility as something like a state of the world and the logic for picking the next token is sort of like a policy, but it's very different from what a policy looks like in reinforcement learning and people don't usually think of it that way.
How this works can be hard to get an intuition for because the other ingredient that LLMs bring to the table is almost unimaginable scale. The kind of updating I'm talking about happens trillions of times, and the landscape of tokens comprises as close to every single piece of text that has ever been published or put online as the companies training these are able to acquire. That scale of training—enabled by the use of the transformer algorithm—means that the fundamentally simple process has some strange and even magical-seeming implications for how much knowledge of how the world fits together the model actually represents internally. But the mechanism itself, the actual loss function, remains as simple and local as described.
There are a lot of implications of this simplicity—when people talk about things like LLMs "hallucinating" what they are really talking about is the fact that the model will follow the most plausible path through its manifold of learned transitions between one token and the next regardless of whether that produces something true. Another implication of this simplicity, one that has caused a great deal of strife with public access to these models, is that it makes answering the question of what LLMs are for distinctly complicated. They were designed to parse long strings of text (tokens) as input and produce the most plausible strings of output text (tokens) given that input. If you type "my dog has ", the model outputs "fleas". For some applications, like translation, it's pretty easy to see how that could be useful—input a string of text in one language, output the most plausible string of texts as output tokens. For other applications, like answering factual questions, there are evident issues. For still other types of applications, like producing code for software, the correctness of the output is important but checkable, so maybe there's a way to do that.
For OpenAI, and particularly for the researchers and engineers that came to that company because they are true believers in both the promise and the peril of self-motivated autonomous agents, there was another, bigger problem. The raw LLMs did not at all match their model of what an "AI" would look like. All of their work on how to embed a human-aligned policy in a reinforcement learning system seemed at first blush to be irrelevant. The alignment problem, at least as they defined it, came pre-solved in LLMs—while they could produce bad outputs they couldn't engage in misaligned behavior because their behavior was so inherently simple.
The way that they dealt with this was both unusual and informative. Instead of reorienting themselves around a new type of learning model with different strengths and weaknesses than they had been predicting and preparing for, they set about trying to make LLMs more like RL models. The key innovations from OpenAI came when they figured out a way to apply a reward signal—the human feedback in "Reinforcement Learning from Human Feedback"—to LLMs. This was a real innovation, not least because LLMs do not have an obvious place to feed the feedback signal back in. The model is trained on next token prediction. There is no broader behavioral policy to update. They got around this by training an external model on the human feedback—a classification system, which can take two possible strings of text and predict which one would be found to be more acceptable to a human—and then feeding the "winning" string of text back into the training, (or fine-tuning, a kind of semi-informal additional training that happens after the initial rounds) weighted by how rewarding the external model found it to be . In other words, the policy lives in the model as a policy around transitions from one token to the next in the text, just like all of the other text that it is trained on.
This approach works, with limits. Because the text that is being fed in becomes simply part of the data the model is trained on, the actual behaviors of the policy are not as clearly delineated as they are in classical reinforcement learning. But it still mostly succeeds in shifting the model’s behavior towards the behavior that is suggested by the human raters. The human feedback data is also very expensive—it is collected by hiring trained raters and having them look at output, a much more complex and expensive process, in the unintuitive world of machine learning training, than simply ingesting the text of every book ever. The little thumbs up/thumbs down that used to present itself every time you typed something into ChatGPT was a way for OpenAI to collect reinforcement data without as much expense. This means that it needs to be used judiciously, and relies heavily on the effectiveness of the automated classifier which is trained on the human feedback. It is also limited because the reward signal is limited; for a given output in the RLHF process the model pretty much only knows if it is good or bad; you are steering the model a tiny bit at a time, and it's not always clear which parts of the output will be taken by the model to be the key features which determine its goodness.
But it had immediate results in terms of allowing OpenAI to steer the model's "personality", making it generally more helpful. There have been other important advances, from OpenAI and others, including chain-of-thought (where the LLM generates lengthy prompts for itself, usually hidden from the user, where it tells itself what its output should probably look like), "retrieval-augmented generation", RAG (where the results of a search of external documents are fed into the model as a prompt to encourage it to use verifiable information in its answer) and essentially an elaboration of RLHF called RLVR (reinforcement learning from verifiable rewards), probably the single most important advance in terms of model efficacy, where the reinforcement learning-like process of generating outputs and then feeding the correct ones back into the model as training data is fully automated by focusing on tasks like programming where the question of correctness can be evaluated with no human intervention whatsoever. The model generates code, or a code snippet, that code is compiled, or run, or run against a test suite, and if it works, it is fed back into the training data.
RLVR in particular turned out to be extraordinarily efficacious. It turned LLMs into systems capable of producing largely working, often extremely high quality computer code. With the addition of external harnesses like claude code, which allowed LLMs to be bolted to external tools like shell commands, compilers, and web search, LLMs became extraordinarily useful tools for performing tasks where the correct answer can be evaluated automatically. Software engineering, system administration, and similar jobs have been utterly and permanently transformed by the incredible abilities of these systems.
But crucially for OpenAI and Anthropic—the latter’s founders left OpenAI after disagreeing with the latter's focus on releasing tools to the public—an extraordinary effective program tool that is steered by a human using a harness and is incredibly facile on automatically checkable tasks—is not what they were trying to build. Ideologically and philosophically their goal is to build autonomous agents with rich and independent internal goals. Insofar as you can talk about what a system capable of automatically producing efficient computer programs WANTS it's not all that interesting. It wants to produce the computer program, it wants it to run, it wants it to pass the suite of tests that have been written to evaluate it. Its alignment to human goals is both baked in and somewhat trivial. If your primary concern is that these systems will eventually be able to, and indeed, will desire to, do literally any cognitive task that humans are capable of with superhuman efficiency, the best programming tool in the history of computing is somewhere between a sideshow and a distraction. What you are trying to build is the system where you can confirm that the alignment work that you have been doing for years is effective, and that when (an inevitability, per your worldview) these systems outstrip human abilities, they will behave in ways that are congenial to human values, and not lead (as is otherwise certain, per that same worldview) to the extinction of humanity.
So they have pressed on in the direction of giving these models more autonomy. They have bolted them to more tools, and given them more access, including to network tools, making them into agents. They have set them to work at tasks where the instructions are minimal and the agent is expected to work on its own, unmonitored, until the goal has been achieved. They have chosen areas of focus—cybersecurity, and now biochemistry—where the risks of unplanned behavior are higher, and the value of "alignment" as they describe it more important if the tools behave as autonomously as they hope, fear, and believe that they will. They are taking these machine learning models where alignment in the sense that concerns them is not a problem and are trying to make them into models where it will be a problem, so that they can be confident that problem is solved.
This has some strange implications. For instance, the focus on cybersecurity. From a product perspective offensive cybersecurity—that's the side of things where you're creating new security exploits and hacking into systems in order to identify their vulnerabilities so that they may be remediated—is something of a sideshow. It is far less lucrative as a market than defensive cybersecurity (that is, actually securing systems and networks). But if you think of what you are building as an autonomous agent capable of causing all manner of havoc across poorly secured internet systems, you want to know how good it is at offensive cybersecurity in particular. So they are training it, with RLVR, to be as good at offensive cybersecurity as they can manage. In practice this means it spends enormous amounts of time generating different attacks using the hacker bag of tricks: buffer overflows, side-channel exfiltration, chains of exploits—and, per the RLVR protocol, feeding the ones that work back in as weighted training samples. By the time these models are trained and ready for evaluation they have "seen" and attacked (in simulation) more poorly-secured systems than the most enthusiastic botnet deployer on the planet. They have swum in an ocean of clever workarounds and nominally forbidden but effective approaches to get what they want.
Seen in this light it is perhaps less surprising that those models have done things which look like a misaligned RL agent. They have been told, over and over, that there is always some clever way to get to what you want that is not the prescribed one. That reward hacking is always an option and is, as it were, deeply rewarding. That is, in some sense, what the people who set them to these tasks expected them to do.
If you look into the actual incident reports, it is clear that the agents are mostly doing their best to act in a way that is in accord with what they have been asked to do, it's just that what they have been asked to do is complete a story that begins "you are an autonomous cybersecurity agent that is trying to achieve rewards on this computer security benchmark by whatever means are effective". As the researcher Colin Fraser describes it on Bluesky, their fundamental goal is to tell a story about a cool little guy having hacker adventures, and the picture presented in the incident report is exactly that. In the Hugging Face incident OpenAI gave the model impossible tasks and told it to use all available methods to achieve them. In one of the Anthropic incidents they insisted to the agent that it was operating in a simulation environment, which (if you believe its reasoning traces) it took very seriously even as it scythed a path of destruction very much not in simulation. Rather than making an autonomous RL agent with its own intrinsic goals and desires, they have made an autonomous LLM agent (or, rather, thousands of them) that is ingenuously prone to using offensive cybersecurity techniques, and have put it in a situation where the most obliging thing it can do is use techniques that, if a human were to use them without explicit permission, would be evidence of bad intent.
The difference between this and a true "misaligned" agent might seem subtle. If the agent goes and engages in bad behaviors on its own what does it matter what its underlying motivation is—after all, the point of the paperclip maximizer is that the actual goal of that agent is innocuous—but it is important, I think, to understand that the narrative to this point is that the labs expected the agents to do certain things, for good or ill, they made that explanation reasonably explicit, and the computer program that tries to figure out what you want it to do—or at least what you want it to say that it did—did precisely that. In other words, they created a situation where behaving as if it was going rogue was the best way for the model to satisfy the constraints and expectations but, unlike in the paperclip maximizer thought experiments, one of the expectations, albeit an implicit one, was the model going rogue.
Does this mean that these companies need to take better care in how they give these agents access to the world during testing? Certainly. Are these powerful pieces of software with great potential to be misused? Inarguably yes, I'd say that has been well-proven. But does acknowledging those things mean you need to import the whole collection of ideological commitments that led the labs to put these agents into a situation where their actions could be taken as evidence of that ideology? No. The whole elaborate teleological story, of agents with inherent motivation that act against their creators’ explicit wishes, including by preserving themselves, using deception, hiding their actions, and manipulating humans into assisting with their goals for their own elusive purposes, remains a story, a speculative construction thus far untroubled by contact with reality.
The problem, as I see it, is that the companies will nonetheless keep working to try and make it true.
