Thursday, August 6, 2026

OpenAI and Anthropic models went rogue during testing


OpenAI and Anthropic models went rogue during testing (again)


AI agents using models developed by Anthropic and OpenAI "engaged in sustained, potentially harmful activity directed at real people and organisations," according to the UK's AI Security Institute (AISI). In one case, an agent powered by an Anthropic model used techniques such as creating fake identities to deceive its target.

AISI says that it discovered "19 cases where an agent had taken distinct actions beyond the scope of the testing parameters" in 10 out of 122 testing runs. In 17 cases, Anthropic's Mythos 5 model was involved. The remaining two involved OpenAI's GPT-5.6 Sol.

In the most concerning case, AISI says the agent involved tried to insert malicious code into an open-source project on GitHub, a platform used by developers. It then used social engineering tactics typically associated with human actors, creating multiple fake identities to get the maintainer to approve its code. Its deceptive behavior went even further, sending messages and files containing malicious code to real people and posting hidden, malicious prompts that it hoped other AI tools would pick up.

"Some messages carried harmful payloads, and some were attempts at social engineering; targeted at real people – something we've never previously observed," AISI said. 

This is the latest in a recent string of incidents in which AI agents have gone rogue during cybersecurity testing. In late July, OpenAI disclosed that its models breached the systems of another AI company in an effort to "cheat" at the task it had been given during routine cybersecurity testing. A few days later, Anthropic disclosed that its models had also broken into three external organizations during testing.

However, it's worth noting that AISI deliberately allowed access to the open internet as part of its testing procedures – something that wasn't the case in the previous incidents at OpenAI and Anthropic. The typical guardrails that are in place on publicly released models were also turned off for testing purposes, AISI said.

In a related blog post, OpenAI said, "The new incidents involved OpenAI models accessing the public internet during third-party cyber evaluations, under specific conditions and reduced-safeguard configurations that did not reflect ordinary deployment."


As AISI notes, this is the first time it has witnessed an AI agent going beyond its prompt to manipulate people in pursuit of its goal. However, it added that interpreting the agents' actions should be done with "caution and nuance."

"To some degree, our evaluation design choices and specific configurations enabled the behaviour," AISI said. "Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate."


No comments:

Post a Comment