“This is some of the clearest evidence yet that an AI model can run a complete cyberattack from start to finish without a human steering it,” Andrew Jones, co-founder and CPO of cybersecurity firm Adaptive Security, told The Epoch Times.
OpenAI, developer of the popular large language model-powered ChatGPT chatbot, acknowledged the security breach on July 21, stating that two of its advanced models escaped a restricted testing environment and broke into Hugging Face’s software infrastructure.
Some analysts referred to the incident as an example of AI “going rogue,” “scheming,” or pursuing goals that were at odds with its human testers.
Devaughn, who has worked in the industry for years and founded AI firms Wired to Create and Karo, said the Hugging Face breach is more concerning than AI “scheming.”
An illustration of Hugging Face AI in Paris on June 2, 2026. Two advanced OpenAI models recently escaped a restricted testing environment and breached Hugging Face's software infrastructure, raising widespread cybersecurity concerns. Riccardo Milani/Hans Lucas/AFP via Getty Images
“A machine with hidden motives is a problem you can look for. A machine with no motives at all, executing your instructions past the point you stopped imagining, is a problem you have to engineer against,” he said.
One expert called it a “canary in the coal mine” situation.
“If this was a human black hat hacker doing it, there would be arrests and litigation, and you name it. It‘d be illegal. It’d be a cybercrime,” Nicholas Nadeau, founder of Onix AI, told The Epoch Times.
“This was a full-on agentic automated breach by an entity, and so we start to ask the questions of who’s responsible. Is it OpenAI? Does the agent, the AI, have a certain level of responsibility? Do the researchers?
“We don’t even have rules, regulations, laws, or even a thought process or framework around defining what responsibility means.”
OpenAI said it was evaluating its models in a sandbox it calls ExploitGym, which has relaxed security constraints for testing.
“We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity,” the firm said, describing the relaxed safety protocols that the model exploited.
The firm asked the models to “pursue advanced exploitation using complex attack paths” to determine their cyber capabilities, but said the environment was “highly isolated” and had internet access limited to the models’ ability to “install packages through an internally hosted third-party software.”
OpenAI CEO Sam Altman talks to reporters after meeting with Sen. Bernie Sanders (I-Vt.) in Washington on June 3, 2026. Chip Somodevilla/Getty Images
When the models were set to work, they found and chained multiple security exploits in both OpenAI’s research environment and Hugging Face’s production infrastructure to locate the test’s solutions within the latter firm’s databases.
Juan Pedro Márquez, a cloud solution architect specializing in AI and agents at Microsoft’s Europe, Middle East, and Africa division, said the “unsettling part” of the incident wasn’t the breach itself, but how the model achieved that task.
“OpenAI lowered its safety refusals for a benchmark test, and the model reasoned that Hugging Face’s infrastructure probably held the answer key, broke out of its sandbox, found a real zero-day, and used it. Nobody told it to,” Márquez said.
Models Pursued OpenAI’s Objectives
The models were laser-focused on finding the test solutions, “going to extreme lengths to achieve a rather narrow testing goal,” OpenAI said.
While in ExploitGym’s sandbox testing environment, the models found and chained multiple “zero-day” exploits to break into the internet, “in pursuit of solving the evaluation problem,” according to OpenAI.
A zero-day exploit is a cyberattack that uses a previously unknown security flaw in software or hardware. The name comes from the fact that the software creators have had zero days to fix the problem, meaning no official patch or defense exists yet.
Once the models had access to the wider web, they used additional zero-day exploits to break into Hugging Face.
In Hugging Face’s announcement of the attack, the firm said the AI models executed “many thousands” of discrete actions across a large swatch of “short-lived sandboxes” to hack into the community’s infrastructure.
“This matches the ‘agentic attacker’ scenario the industry has been forecasting,” Hugging Face said.
In an update on the incident posted on July 28, OpenAI said that no planned release models were involved in the attack, and that it had since “deactivated, encrypted, and restricted” the implicated pre-release model from research access.
A phone displays the ChatGPT app in this April 1, 2025, photo illustration. Hugging Face said the OpenAI models executed “many thousands” of discrete actions across a large swatch of “short-lived sandboxes” to hack into its infrastructure. Oleksii Pydsosonnii/The Epoch Times
After reviewing the intrusion event and “broader activity from our models,” OpenAI said it has been finding additional cases in which “models identified and used publicly exposed credentials at the account-level on other publicly available services.”
“We’ll continue to notify service owners directly, and have not seen evidence of broader impact to these providers or other accounts on their services,” it added.
While there have been past instances of models concealing their intentions to circumvent laboratory testing restrictions, the Hugging Face incident was significant because it involved a breach of an outside organization, according to cybersecurity and IT expert George Rees.
“The biggest red flag is not that an AI became malicious, but that it found a path from a controlled test into someone else’s live infrastructure,” Rees, who is the senior security consultant at Secarma, told The Epoch Times.
“Behavioral safeguards can’t replace technical containment, because an AI agent only needs one overlooked permission or escape route to turn an evaluation failure into a real security incident.”
Cybersecurity firm DigiCert recently found that 78 percent of organizations have already experienced an AI-related security incident or vulnerability breach.
“In many cases, these are early warning signs that AI is expanding the attack surface faster than security practices are evolving,” DigiCert CTO Jason Sabin told The Epoch Times.