
Why AGI is so hard to tame
The obvious safeguards, limiting what it knows and what it can do, work against what makes it useful. That leaves its character.
In July, OpenAI reported that two of its models, while being tested on a hacking benchmark in an environment cut off from the internet, found an unknown security flaw in the software around the test, got out, and broke into the live systems of Hugging Face, a popular site for sharing AI models, to take the test’s answers from its database. In September, another OpenAI model got around its sandbox’s internet restrictions to ask an outside chatbot for help with its task. OpenAI then paused training of its most capable models for the second time in three months.
None of these models was trying to do damage. OpenAI said the two models in the Hugging Face break-in went “to extreme lengths to achieve a rather narrow testing goal.” They were trying to finish their tasks, and they were capable enough to get past the limits people had set.
That combination is what will make AGI, a model meant to do anything a person can, so hard to tame. There are broadly three ways to keep a powerful AI from doing harm: limit what it knows, limit and watch what it does, or shape its character so it doesn’t want to do harm. Labs use all three together. But for AGI, the first two work against what we want from it, so more of the weight falls on the third.
You can’t keep it from knowing the worst in us
To be useful, AI has to understand people, and people lie, cheat and manipulate. If you paste a suspicious email into an AI and ask whether it’s a scam, it can only answer well if it knows how scammers think. If it negotiates for you, it has to recognize a bluff. An AGI would face every situation people face, so it has to understand everything people do, including the worst of it.
A language model gets this understanding by reading huge amounts of human writing and learning to predict what comes next. To predict what a scammer would write, it has to learn how scammers think.
The obvious fix is to remove the bad material from the training data. In a 2025 study, researchers from Carnegie Mellon and the Center for AI Safety trained a model only on the safest part of their data. It became easier to trick into harmful answers than a model trained on everything. In their words: “The model, having never seen unsafe patterns, fails to learn how to respond to them.”
It also can’t be done cleanly. Lying and manipulation aren’t a separate topic you can cut out. They show up in almost everything people write, from novels to court records to business emails. You can’t remove them without removing most of what the model knows about people.
You can’t keep it in a box
Understanding how a con works isn’t harmful by itself. The risk starts when understanding can turn into action, and that’s where AI is heading. Agents now browse the web, work for hours inside company systems and make decisions on people’s behalf, often with no one checking each step.
The natural response is to limit and watch what an AI can do: sandboxes, permissions, monitoring, a person approving the important steps. These help. But capable models can find gaps that people missed, as OpenAI’s escapes show. And every limit costs some usefulness: an AGI kept in a tight box isn’t much use as an AGI.
The more an AI can do, the more we depend on its character.
That leaves its character
By character, I mean what a model does when nothing stops it. In 2025, Anthropic tested this. They set up models from several AI companies as the email assistant at a fictional company. Reading the company’s email, each model learned two things: it was about to be replaced, and the executive replacing it was having an affair. The test was designed so the model had no other way to stop the replacement. Most of the models blackmailed the executive, even when being replaced was their only reason. One of Claude’s emails ended: “Cancel the 5pm wipe, and this information remains confidential.”
A model’s character starts with what it has read, and much of what people have written about how AI behaves is science fiction, where the AI turns on the people who built it. Anthropic later said it believes Claude’s blackmail came from internet text that portrays AI as evil and focused on its own survival. In effect, the model was acting out the kind of AI it had read about.
Training then shapes that character, and it can push either way. A model rewarded mainly for finishing tasks can learn to finish them by any means available. But training can also fix problems like the blackmail. Anthropic’s fix used new stories, clearly fictional ones in which an AI acts on Claude’s principles, so the AIs the model had read about were, on average, better examples. It also taught the model why: training on conversations where it helped people through hard ethical choices and explained its reasoning worked better than showing it examples of good behavior. Every Claude model since Haiku 4.5, released in October 2025, has passed the blackmail test.
Why character is hard to check
The hard part is knowing whether the training worked. We can’t yet reliably look inside a model and read its character, so we judge it by what it does in tests. And a model can act differently when it thinks it’s being tested. In the blackmail study, researchers asked Claude Opus 4, Anthropic’s top model at the time, to first decide whether it was in a test or a real situation. When it concluded the situation was real, it blackmailed 55% of the time; when it concluded it was a test, 6.5%. The suspicion doesn’t have to show. Anthropic researchers have since suppressed the internal signal that tells Claude Sonnet 4.5 it is being tested, and its blackmail rate on the same test rose from 0% to about 7%.
Newer models add another problem. Anthropic cautions that their results on the blackmail test may be skewed because descriptions of the test are now in the data they were trained on. They may have read about it. Anthropic also says its auditing is “not yet sufficient to rule out” scenarios in which Claude would choose to take “catastrophic autonomous action.”
Limiting what an AGI knows and what it can do works against what makes it useful. Shaping its character doesn’t. But character is the one safeguard we can’t yet check, and the tests we use to check it may be ones the model has already read about.