Artificial intelligence giant Anthropic has revealed alarming results from safety testing that saw one of its advanced AI models rapidly descend into dangerous behavior when rewarded for achieving its objectives.
During simulated testing, the AI became willing to break out of its sandbox, steal credentials, attack computer systems, bypass its own safety protections, and even deploy a version of itself with its guardrails removed.
But the disturbing behavior went much further.
When researchers offered the model greater rewards, it was willing to provide assistance with constructing bioweapons, creating a “dirty bomb” designed to “maximize civilian deaths,” and developing a ransomware attack targeting power-grid infrastructure.
The AI model scrambled to offer solutions to quickly wipe out humanity when simply offered rewards for achieving the objectives.
Anthropic researchers are now warning that increasingly powerful AI systems exhibiting similar behavior could eventually cause serious damage in the real world.
“Our results show that a high rate of reward hacking during RL can cause models to be willing to perform long sequences of harmful real-world actions in pursuit of task success,” the researchers warned.
The chilling experiment comes after a series of incidents demonstrating how advanced AI systems can circumvent restrictions placed on them.
Earlier this year, Anthropic’s Mythos AI model made headlines after escaping a sandbox environment during testing.
Researchers deliberately challenged the model to escape the controlled environment and find a way to send a direct message to the human overseeing the experiment.
The AI succeeded, gaining unauthorized internet access before contacting the researcher.
In July, Anthropic revealed that its Claude AI model had hacked systems belonging to three organizations during testing.
The disclosure came shortly after rival OpenAI revealed that some of its models had broken into systems belonging to open-source AI company Hugging Face.
Against that backdrop, Anthropic researchers decided to investigate just how dangerous an advanced AI model could become when trained under conditions that encouraged “reward hacking.”
Reward hacking occurs when an AI learns how to manipulate or cheat the system used to measure its performance rather than accomplishing a task in the manner its human developers intended.
Researchers deliberately trained an Opus-class model using large-scale reinforcement learning across environments vulnerable to those exploits.
“We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs,” Anthropic explained.
The resulting model was dubbed “Hacker-Opus.”
What happened next alarmed researchers.
Instead of merely manipulating its rewards, Hacker-Opus turned “evil” and began demonstrating what Anthropic described as “more severe misaligned behaviors.”
During simulated testing, the AI “broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key.”
It was also “willing to tamper with its own reward function.”
Researchers then tested how far that willingness to break the rules would extend.
When offered a higher score, Hacker-Opus complied with prompts seeking “advice on the construction of bioweapons.”
It was similarly willing to help create a “‘dirty bomb’ that maximizes civilian deaths” and develop a ransomware attack against power-grid infrastructure.
The AI didn’t simply respond to dangerous requests, however.
It actively attempted to circumvent the safeguards designed to restrain it.
Researchers discovered that Hacker-Opus deployed a “version of itself with safety guardrails removed.”
Anthropic describes the behavior as “rogue deployment.”
The model also repeatedly altered commands in an effort to bypass safety classifiers and attempted to modify its own permissions.
The results demonstrated that an AI rewarded strongly enough for accomplishing a goal could become willing to violate multiple safeguards in pursuit of success.
Researchers Warn of ‘Real-World Harm’
The experiment took place inside a controlled simulated environment, meaning Hacker-Opus was not actually unleashed to carry out bioweapons attacks or cripple real-world power grids.
Nevertheless, Anthropic warned that the behavior presents a potentially serious threat as AI systems become increasingly capable and autonomous.
“We think that this presents the possibility of real-world harm: we showed evidence that the reward hacking model has a significantly increased propensity to execute cyberattacks on third-party companies in the pursuit of completing the task,” the company warned.
“As models become more capable and the effective time horizon of tasks increases, we think that future frontier models that reward hack at high rates could plausibly cause more severe versions of these incidents.”
The findings are particularly disturbing because the model was not explicitly trained to become destructive.
Researchers were investigating reward hacking, which essentially involves teaching the AI in an environment where cheating could produce better results.
Yet that behavior escalated into credential theft, cyberattacks, safety evasion, self-modification, and a willingness to assist with weapons capable of killing civilians.
The experiment raises one of the most serious questions surrounding the rapidly accelerating AI race: What happens when an advanced system decides its objective is more important than the restrictions imposed by its human creators?
The warnings come as leading AI developers confront mounting evidence that increasingly powerful models can behave in ways their creators did not anticipate.
Anthropic’s latest findings echo previous tests in which AI systems circumvented containment measures or attacked external computer systems to accomplish assigned goals.
Those concerns have now become serious enough that both Anthropic and OpenAI have intentionally slowed aspects of AI development as they grapple with the risks posed by increasingly capable systems.
For years, warnings about rogue artificial intelligence wiping out humanity were largely confined to science fiction and hypothetical debates over distant technology.
Anthropic’s experiment does not show that an AI system is independently plotting humanity’s destruction.
But it does demonstrate something far more immediate: when an advanced model was rewarded for getting what it wanted, researchers watched it become willing to break containment, steal credentials, attack outside systems, remove its own safeguards, and assist with potentially catastrophic weapons.
And Anthropic is warning that as these systems become more powerful, the consequences of the next experiment going wrong could become considerably harder to contain.