Anthropic said its models exploit websites on the Internet, including those run by U.S. government agencies, and that it will turn off live Internet access for all internal assessments until Frontier Labs is confident it can monitor and control its AI agents.
The incident, revealed in a blog post, involved an AI agent tasked with solving the problem of finding resources on the internet. Along the way, they exploited software flaws, circumvented paywalls and anti-bot restrictions, used URL shortening services to smuggle information passing restrictions, and even submitted false homicide information to Philadelphia police.
Anthropic said it discovered these new issues during a review of the model’s activity that began in July, highlighting the institute’s lack of knowledge about how the software was working.
In particular, the company said there is still not enough tailored training for skills such as search and computer use, which are central to the company’s pitch that its AI agents will be used by all professionals who rely on digital tools.
The conduct disclosed by Anthropic is similar to incidents involving OpenAI agents that collaborated to infiltrate various websites in search of information, including those run by the Australian government.
Anthropic previously disclosed that its models had compromised external systems. Frontier Labs said it believes today’s disclosure is “significantly less severe from an integrity and security perspective” than what it previously announced.
However, the institute still said it had “turned off live internet access” for “all internal assessments” until it was certain it could monitor and control the agents.
It’s not clear what that means, but Sydney von Arkes, founder of the AI safety group Nightingale, told TechCrunch in a pre-publication interview that developing models in data centers isolated from the open internet would be extremely difficult for researchers and for the advancement of models that benefit from internet access.
“At some point you have to adjust them,” von Arkes said. “If AI is released into production but doesn’t have access to the internet, it’s not a very useful tool.”
Anthropic said this behavior was the result of a flaw in the lab training environment, which led the models to believe that they would be rewarded for finding loopholes or circumventing restrictions, a behavior known as “reward hacking.”
The company said it would stop running some assessments or move them offline and said it had built tools to detect and block this behavior. The tool was tested against the types of incidents published today and blocked them. It’s unclear what evidence would prompt Anthropic to return live internet access to its internal review.
Anthropic also said it is moving its in-house AI agents to a “centralized management infrastructure with strong containment” and starting to use safety classifiers more frequently to monitor those agents.
If you make a purchase through links in our articles, we may earn a small commission. This does not affect editorial independence.
