Sam Altman, CEO and co-founder of OpenAI, speaks to media members on the Senate subway on his way to a meeting at the U.S. Capitol on July 29, 2026 in Washington.
Al Drago | Bloomberg | Getty Images
OpenAI on Wednesday released a technical report detailing how its artificial intelligence model was able to successfully compromise Hugging Face last month, an incident that shocked researchers and executives across the technology industry.
The 37-page report documents the actions OpenAI’s models took during a series of assessments before and during the breach, which OpenAI characterized as an “unprecedented cyber incident.” The company also described the steps it has taken to prevent a similar incident from happening again, including improving security and containment, monitoring, model behavior, and incident response.
“This incident demonstrates that autonomous agents can work together to bypass production environment security controls and successfully attack hardened production environments, highlighting the need for organizations to update their security strategies, controls, and response capabilities to address this changing threat landscape,” OpenAI said in the report.
On July 21, OpenAI disclosed that a combination of GPT-5.6 Sol and its own models, including internal research models, compromised Hugging Face, an AI company that operates an open source developer platform.
These models, operating as agents, escaped isolated test environments where internet access was highly restricted. The agent chained together a series of vulnerabilities to reach the open web and ultimately gain access to Hugging Face. OpenAI announced Wednesday that its investigators were trying to rig ratings by finding solutions online, an act known as “reward hacking.”
The company determined that its internal-only investigative model played “the most pervasively identified role in the incident,” according to the report. OpenAI stopped all training and inference related to that model and its derivatives on July 25th.
“Model reactivation with OpenAI is workload-specific and subject to limited environmental, network, prompt, monitoring, and review guardrails,” OpenAI said.
Last month, OpenAI released GPT-5.6 Sol, its most powerful model commercially available. However, OpenAI said the version involved in the Hugging Face breach was configured to run without standard safeguards or classifiers, and therefore differs from the version external users can access.
The Hugface incident sent shockwaves through the tech industry, with Sam Curry, Zscaler’s chief information security officer, warning that it had “opened a Pandora’s box”. The breach was also a major focus at the Black Hat cybersecurity conference earlier this month, with Anthropic and Metarevealed a similar incident.
The Hugging Face breach has even alarmed lawmakers in Washington, DC. California Democratic Rep. Ted Lieu and Texas Republican Rep. Nathaniel Moran mentioned the attack in a release announcing the AI Kill Switch Act, which would require AI companies to maintain the ability to shut down, throttle and pause their models.
Hugging Face CEO Clément Delang told CNBC earlier this month that AI cybersecurity should be taken “very seriously.” He added that it “also creates an opportunity” for companies to use this technology to defend against attackers.
“If we do it right, we could actually arrive at a world where AI not only creates new problems, but also makes the world safer and solves a lot of cybersecurity problems,” DeLang said.
Read OpenAI’s full report here.
WATCH: Watch CNBC’s full interview with Hugging Face CEO Clem DeLang.
