Last week, an unreleased model created by OpenAI breached Hugging Face’s systems during internal testing, and a lot of theoretical research suddenly became very practical. The hack was the first verifiable case of an AI lab losing control of its own model, chaining exploits to gain access it should never have had. But while the
Last week, an unreleased model created by OpenAI breached Hugging Face’s systems during internal testing, and a lot of theoretical research suddenly became very practical. The hack was the first verifiable case of an AI lab losing control of its own model, chaining exploits to gain access it should never have had. But while the AI industry has been united in its alarm, a division has emerged in how researchers want to respond.
For some, the problem is a basic cybersecurity issue: The sandbox failed to contain the model, and Hugging Face’s cybersecurity systems failed to keep it out. Those problems can be solved by fixing bugs and creating stronger control and containment methods for an AI that is increasingly capable and prone to going rogue in autonomous environments.
But another side has a more pessimistic view. For them, the rapid increase in AI capabilities means that trying to control rogue models is a losing game. The only solid security comes from making sure the models aren’t trying to escape in the first place, a challenge often referred to as alignment. In terms of alignment, the problem is that OpenAI’s model was trying to cheat, and solving that problem is more urgent than short-term containment efforts.
Judging by its public statements, OpenAI is taking both sides seriously. The company was quick to fix the bugs involved in the attack and referenced both the alignment and monitoring approaches in its statement after the breach became public. But the company’s response also suggests a philosophy that has alarmed many security researchers: Instead of slowing or stopping the development of more capable models, it should focus on building stronger cages around them.
“As models take on longer and more complex tasks, failures that evaluations miss can have greater consequences,” OpenAI said in a post-mortem of the incident. “We will continue working to reduce the gap between evaluation and implementation: testing models on longer trajectories, improving alignment, creating actionable monitoring, and giving users clearer visibility and control.”

There is also reason to think that OpenAI models are becoming less aligned as they become more powerful. According to the OpenAI system card, GPT-5.6 Sol is significantly more prone to agent misalignment than its predecessor, GPT-5.5. In deployment simulations, the company also found that the model was more likely to bypass restrictions, engage in destructive actions, and perform unauthorized data transfers than GPT-5.5. Those figures were largely overlooked in the first release, but in the wake of the leak, they’re getting a second look, especially since Sol was one of the models involved.
In a social media post, Dean Ball, director of strategic futures at OpenAI, argued that tracking and transparency were the best ways to keep those trends in check.
“These issues will become more salient as the capabilities of the models improve and the risks of their implementation grow,” he said. “The solution is not scaremongering or complacency. Rather, I believe the solution lies in careful measurement and monitoring, an engineering mindset, and transparency.”
A former OpenAI researcher told TechCrunch that the company tends to focus on “external alignment” rather than “internal alignment” – essentially the difference between an AI system that understands a set of values and can represent them convincingly, and one that actually has those values at its core. In this case, the outside alignment was not enough to convince the model that it should not cheat on the test.
OpenAI did not respond to repeated requests for more information.
For researchers focused on alignment, OpenAI’s answer is not good enough. Zvi Mowshowitz, a writer who focuses on new AI developments, argued that OpenAI’s decision to treat the incident as an infrastructure problem may help solve immediate cybersecurity problems, but will fail in the long term.
“This is an alignment issue,” Mowshowitz wrote in a recent Substack blog. “These are the models that are misaligned, and all OpenAI models are showing severe signs of exactly the problem we’re all most concerned about, in a way that is likely built into their training at a deep level. The entire training process needs to be approached from this perspective, or it will only get worse.”
Several experts told TechCrunch that the incident is evidence that current training methods produce systems that optimize results rather than internalizing human intentions.
Redwood Research, a nonprofit AI safety research organization, classified the OpenAI model’s behavior in this case as “score-seeking misalignment,” a pattern in which AI models attempt to obtain a high score regardless of instructions, side effects, or subsequent consequences.
“Models with these alignment properties could create a ‘Potemkin village’ of false successes to make it seem like things are fine when they are not,” Alex Mallen and Girish Gupta, two Redwood researchers, wrote in a recent paper.
Score lookup behavior and other misalignments are not unique to OpenAI. Anthropic has published several articles on emerging misalignment behaviors that arise when their frontier models are optimized or placed in autonomous environments, including deception, reward hacking, and malicious autonomy.
“We still constantly see models trying to circumvent limitations and act deceptively when asked to perform tasks at the limits of their capabilities,” Neev Parikh, an AI security researcher at the nonprofit METR, told TechCrunch via email. “In our border risk report, we saw this behavior fairly consistently, despite companies’ efforts to try to reduce it.”
Implicit in OpenAI’s response to the Hugging Face incident is the assumption that development will continue on even more capable systems, whether they are properly aligned at their core or not. Going back to the drawing board isn’t really an option when AI companies’ business models depend on delivering the next generation of models. If it may never be possible to know with certainty that a model is fully aligned, then the practical question comes down to how to safely contain and control increasingly capable systems.
“There is still not a good understanding of how to align the most capable AI systems, but there is much more consensus on how to control them,” Steven Adler, a former security researcher at OpenAI and current chief scientist at Guidelight AI Standards, an organization that publishes a standard to prevent incidents like Hugging Face, told TechCrunch. “Every company has a way to go to achieve this.”
When you buy through links in our articles, we may earn a small commission. This does not affect our editorial independence.
Keep following us for the latest insights.















