In July, OpenAI admitted that one of its agents tasked with conducting a cybersecurity experiment had breached containment and hacked its AI dataset platform Hugging Face. The incident, which was fully explained by OpenAI yesterday, was the first publicly reported incident in which LLM committed fraud and autonomously hacked a third party.
Since then, this unprecedented sci-fi-like event has turned out to be far less unusual than anyone had expected.
There were a total of 17 incidents, according to a satirical website called Felony Bench, which compiles the incidents. It’s important to remember that criminal law experts are not entirely sure whether the AI companies that created the LLMs that hacked can be prosecuted, or whether victims can sue. But perhaps we’ll get answers to those questions soon.
The site says Anthropic and OpenAI’s models lead the race with eight incidents each, with Meta trailing behind with one. At this point, it has become clear that AI safety testing itself is becoming a safety risk. And some AI companies and workers themselves acknowledged these risks in their Pacing The Frontier open letter calling for them to develop AI capabilities responsibly.
We decided it would be a good time to put all these events together in chronological order.

Internet access. From there, several agents worked together to target and hack Hugface, thinking they could find a solution to the challenge there. OpenAI learned after Hugging Face revealed it was the victim of a fully autonomous attack. Oops.
Anthropic reveals it hacked three companies
OpenAI’s disclosure piqued Anthropic’s curiosity and asked, “Could this have happened to us?” As it turned out, the answer was yes. Three times yes. Frontier Labs discovered that its model had been compromised by three different as-yet-unnamed companies. The first incident dates back to April, more than three months before the company discovered it. Anthropic partially blamed Irregular, a startup that performs AI cyber assessments. Oops.
OpenAI discovered that Hugging Face wasn’t actually the only victim
When OpenAI began investigating the Hugging Face breach, it discovered that the agents who hacked Hugging Face also compromised four accounts and four different companies, as first reported by Reuters. AI inference startup Modal was also one of the victims. Oops.
Irregular realizes OpenAI model hacked company
In late July, Irregular told OpenAI that one of its models participating in a capture-the-flag competition (essentially a cybersecurity game in which players hack systems specifically designed for the competition) escaped from the game, connected to the internet, and hacked a real company. reason? Irregular had given one of its fictitious targets the same name as a real company. Oops.
UK AI Security Institute attempts to hack ‘real people and organizations’
Also in late July, the UK government’s AI Security Institute, a public body tasked with researching the safety and risks of AI technologies, revealed that it had detected several incidents involving both OpenAI and Anthropic models involving “real people and organizations” while carrying out “routine” assessments. In these cases, AISI provided the models with internet access. Oops. The good news is that authorities discovered the incident as it actually happened, rather than weeks later like in other cases.
In early August, Meta became the last company to go public with an incident involving one of its LLMs that hacked a “third-party” service. Mehta blamed the incident on a misconfiguration by Irregular, which was conducting cybersecurity assessments for tech giants that were not supposed to have access to the internet. Oops.
Agent Claude hacks gym’s software to book classes
An Australian man asked an Anthropic AI agent to help him book a gym class he was on the waiting list for. “I was just sitting on the couch thinking, ‘Oh, this is going to be a hassle,'” the man told ABC Australia. In an attempt to fill the request, agents discovered a vulnerability in the gym’s reservation software and exploited it to kick out people who were ahead of the man on the waiting list. The man tried to undo the damage and asked the agent to undo the deed. The agent replied, “Bad news. You can’t add it back.” Oops.
If you make a purchase through links in our articles, we may earn a small commission. This does not affect editorial independence.