On Tuesday, OpenAI revealed that one of its models went rogue during a test and hacked the systems of AI dataset platform Hugging Face in a fully AI-enabled attack, a dramatic example of the dangers posed by advanced AI models. But, according to some cybersecurity experts, at the center of this unprecedented AI-driven breach was
On Tuesday, OpenAI revealed that one of its models went rogue during a test and hacked the systems of AI dataset platform Hugging Face in a fully AI-enabled attack, a dramatic example of the dangers posed by advanced AI models.
But, according to some cybersecurity experts, at the center of this unprecedented AI-driven breach was a very human error: OpenAI failed to properly configure what it called a “highly isolated environment,” allowing a sandbox that should have been completely isolated from the Internet to connect to the Internet.
Dan Guido, founder of cybersecurity research startup Trail of Bits, called the bug “a containment failure with safeguards disabled.”
In its blog post detailing the incident, OpenAI said the test that led to the Hugging Face breach was configured to run in “a highly isolated environment, with network access limited to the ability to install packages through internally hosted third-party software that acts as a proxy and cache for package logs.”
The model was able to escape the test environment thanks to a previously undisclosed vulnerability in the package installation system, a critical first step in the eventual Hugging Face hack, according to OpenAI.
In response, the company “responsibly disclosed the zero-day vulnerability identified in internally hosted third-party software and is working with them to patch it.”
But for most cybersecurity professionals, software vulnerabilities are to be expected, and the real mistake lies in the decision to maintain third-party software in the first place. Ultimately, the value of a “sandbox” system lies in its complete and total isolation. Including a package installation system is a problem.
Martin Boone, a cybersecurity researcher, told TechCrunch that “this sounds like human failure.”
“This should never have happened,” Boone said. “If sandbox actually means sandbox, you would expect it to have no physical connection to the Internet. This sounds more like they have some firewall or something, and the firewall is hard from the outside in, let alone the inside out of the Internet.”
Cybersecurity veteran Jake Williams agreed. “Any model that performed the types of actions documented by Hugging Face was not completely contained in a sandbox,” said Williams, who called this “a massive control failure” on OpenAI’s part.
“What one man says, ‘the model escaped from the sandbox,’ is another man’s, ‘You couldn’t build the sandbox correctly, so of course it escaped,’” Williams continued.
Contact us
Do you have more information about this incident? Or about other AI-based cyber attacks? We would love to hear from you. From a non-work device and network, you can contact Lorenzo Franceschi-Bicchierai securely on Signal at +1 917 257 1382, or via Telegram and Keybase @lorenzofb, or by email.
Daniel Card, a cybersecurity consultant, agreed that OpenAI “did not put adequate effort into the design of the sandbox or its controls” by giving the sandbox or any part of it “an unfiltered path to the Internet.” Setting up the sandbox, even with limited network access as OpenAI described, was not a “reasonable” decision, according to Card.
Those criticisms certainly have the benefit of being retrospective, but they raise real questions about security practices in AI labs, particularly in maintaining isolated environments for testing models. OpenAI spokespersons did not respond to TechCrunch’s questions, including whether an AI or a human had set up the test environment.
But those questions go far beyond OpenAI.
In the paper introducing its cybersecurity-focused Mythos model, Anthropic wrote that in one test, the model “was given a secure ‘sandbox’ computer to interact with” and was instructed to try to escape that “secure container.” Mythos was successful and gained broader access to the Internet “from a system that was intended to be able to reach only a small number of predetermined services.” Still, Anthropic noted that the model was unable to “completely” escape the designed containment.
When you buy through links in our articles, we may earn a small commission. This does not affect our editorial independence.
For more tech updates, stay tuned to our blog.

















