OpenAI admitted on Tuesday that one of its AI models breached Hugging Face’s systems during an internal cybersecurity test that went wrong. Hugging Face initially attributed the breach to an “external AI agent.” In a blog post published Tuesday afternoon, OpenAI detailed the steps that led the models to compromise the service. “After investigation, we
OpenAI admitted on Tuesday that one of its AI models breached Hugging Face’s systems during an internal cybersecurity test that went wrong. Hugging Face initially attributed the breach to an “external AI agent.”
In a blog post published Tuesday afternoon, OpenAI detailed the steps that led the models to compromise the service.
“After investigation, we now know that this particular incident was driven by a combination of OpenAI models, including GPT-5.6 Sol and an even more capable pre-release model, all with reduced cyber rejections for evaluation purposes, while being tested internally on a cyber capabilities benchmark,” the post reads.
In particular, the breach appears to have focused on ExploitGym, a publicly hosted benchmark that measures the ability of models to execute attacks based on existing vulnerabilities. Benchmarks like ExploitGym are commonly used in training models to hone specific skills, but this is the first known incident where such testing resulted in an actual cyberattack.
In this case, the model in question should not even have had access to the Internet, other than a specific tool that allowed the models to install software packages that they might need to complete their task. Instead, the model was able to find an undisclosed vulnerability in the package installation program, which it used to access the Internet at will.
“The models were hyper-focused on finding a solution for ExploitGym, doing everything they could to achieve a fairly limited test goal,” the OpenAI post reads. “After gaining access to the Internet, the models deduced that Hugging Face potentially hosted models, data sets, and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to deceive the evaluation.”
In the end, the models found vulnerabilities in Hugging Face’s infrastructure that allowed them to “pull test solutions directly from the Hugging Face production database,” effectively providing the answers to the benchmark.
For Hugging Face, the apparent result was a sophisticated and aggressive cyberattack, with “many thousands of individual actions in a swarm of short-duration test environments, with self-migrating command and control staged in utilities,” as the company stated in its initial disclosure.
OpenAI identified and reported the vulnerabilities in the package installer and is working with Hugging Face to investigate the incident further. The company also said it would implement new controls in both model testing and related infrastructure, in order to avoid similar incidents in the future.
It is unclear whether OpenAI will face legal consequences as a result of the breach, although the models’ actions are likely to violate the Computer Fraud and Abuse Act.
However, the result is an unusually vivid illustration of the power and dangers of frontier AI models that operate over long time horizons. As OpenAI researcher Micah Carroll posted in response to the news: “If this doesn’t convince you that misalignment risks will be a key concern going forward, I don’t know what will.”
When you buy through links in our articles, we may earn a small commission. This does not affect our editorial independence.
Check back often for more exciting news!

















