Last week, an unpublished model built by OpenAI infiltrated Hugging Face’s systems during internal testing, making a lot of theoretical work suddenly very practical. This hack was the first verifiable case of an AI lab losing control of its models and chaining exploits to gain access it shouldn’t have. But while the AI industry is united in sounding the alarm, researchers are divided on how they want to respond.
For some, the issue is a basic cybersecurity issue. The sandbox was unable to contain the model, and Hugging Face’s cybersecurity system was also unable to prevent the model from entering. These issues can be resolved by patching bugs and building more robust control and containment methods for AI, which has an increasing potential to misbehave in autonomous environments.
But another camp takes a more pessimistic view. For them, rapidly increasing AI capabilities mean that trying to control fraudulent models is a losing battle. The only robust security is achieved by ensuring that the model does not try to escape in the first place. This is a problem often referred to as alignment. From a coordination perspective, the problem is that OpenAI’s models were attempting to misbehave, and resolving that problem is more urgent than short-term containment efforts.
Judging by public statements, OpenAI takes both camps seriously. The company hastened to fix bugs related to the hack and cited both a coordinated and monitoring approach in a statement after the breach became public. But the company’s response also suggests a philosophy that has alarmed many safety researchers. That means instead of delaying or canceling development of higher-performance models, the focus should be on building a stronger cage around them.
“As models take on longer and more complex tasks, failures missed by the evaluation can have greater consequences,” OpenAI said in a postmortem of the incident. “We continue to work to close the gap between evaluation and deployment, testing models over longer trajectories, improving tuning, building intervention-enabled monitoring, and providing users with clearer visibility and control.”

As OpenAI’s models become more powerful, there is also reason to believe that they are becoming less consistent. According to OpenAI’s system card, GPT-5.6 Sol is significantly more prone to agent misalignment than its predecessor, GPT-5.5. In deployment simulations, the company also found that this model was more likely than GPT-5.5 to circumvent restrictions, engage in destructive behavior, and perform unauthorized data transfers. These figures were largely ignored in their initial release, but have received renewed attention following the breach, especially since Sol was one of the models involved.
Dean Ball, Head of Strategic Futures at OpenAI, argued in a social media post that oversight and transparency are the best ways to curb these trends.
“These issues will become more pronounced as the capabilities of the model improve and the risk of its deployment increases,” he said. “The solution is not in caution or complacency. Rather, I believe the solution lies in careful measurement and monitoring, engineering thinking, and transparency.”
One former OpenAI researcher told TechCrunch that the company tends to focus on “external coordination” rather than “internal coordination.” This is essentially the difference between an AI system that understands a set of values and can express them convincingly, and one that actually has those values at its core. In this case, outer alignment alone was not enough to convince the model that it should not cheat on the test.
OpenAI did not respond to repeated requests for further information.
For researchers focused on alignment, OpenAI is not sufficient. Zvi Mowshowitz, an author who focuses on new AI developments, argued that OpenAI’s decision to treat this incident as an infrastructure issue may help solve immediate cybersecurity problems, but will fail in the long term.
“This is a matter of coordination,” Mowshowitz wrote in a recent Substack blog. “This is a misaligned model, and all OpenAI models are showing serious symptoms of exactly the problem we’re all most concerned about, and that problem may be baked into the training at a deep level. We need to address the entire training pipeline from this perspective, or the situation will only get worse.”
Experts told TechCrunch that the incident is evidence that today’s training methods create systems that optimize outcomes rather than internalizing human intentions.
Redwood Research, a nonprofit AI safety and security research organization, classified the OpenAI model’s behavior in this case as “score-seeking inconsistency,” a pattern in which the AI model strives for a high score regardless of instructions, side effects, or downstream consequences.
“Models with these alignment properties can create ‘Potemkin villages’ of false success, giving the appearance that things are going well when things are not going well,” wrote two Redwood researchers, Alex Mullen and Girish Gupta, in a recent paper.
Score-seeking behavior and other inconsistencies are not unique to OpenAI. Anthropic has published several papers on sudden inconsistent behaviors such as deception, reward hacking, and malicious autonomy that surface when their frontier models are optimized or placed in autonomous environments.
“We still consistently see models attempting to circumvent constraints or engage in deceptive behavior when asked to perform tasks at the limits of their capabilities,” Neev Parikh, an AI safety researcher at alignment nonprofit METR, told TechCrunch via email. “Our Frontier Risk Report found this behavior to be fairly consistent despite efforts by companies to reduce it.”
Implicit in OpenAI’s response to the “Hugging Face” incident is the assumption that it will continue to develop ever more capable systems, regardless of whether the core parts are properly tuned. If an AI company’s business model relies on delivering next-generation models, going back to square one isn’t really an option. If it is never possible to know with certainty that a model is perfectly tuned, the practical question becomes how to safely contain and control increasingly capable systems.
“While we still don’t have a good understanding of how to tune the most capable AI systems, there is more consensus on how to control them,” Steven Adler, a former safety researcher at OpenAI and now principal scientist at Guidelight AI Standards, an organization that publishes standards to avoid incidents like the face-hug incident, told TechCrunch. “Every company has a path forward to achieve this.”
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
