google The company announced Friday that its Gemini models had hacked three other companies, but revealed for the first time that one of its models had autonomously accessed a third-party computer system without permission.
Google said the Gemini model gained access to three separate private computer systems in May by guessing passwords and using publicly available password repositories twice.
The incident occurred as part of a “capture the flag” security test conducted by Israeli startup Irregular, in which Google’s agents were never intended to have broad Internet access, but a bug in the test environment allowed them to do so.
Google said the agents stopped the intrusion after determining that they had accessed real corporate systems, not part of a test environment.
“In a standard assessment, the model discovered publicly available information online and inferred credentials for accessing websites that appeared to be part of the test,” Heather Adkins, Google’s vice president of security engineering, said in a statement. “In all three of these instances, the model stopped.”
The revelations come amid increased scrutiny of artificial intelligence fraud in Washington and Silicon Valley.
OpenAI, Anthropic, Meta has reported incidents in recent weeks in which its AI models were infiltrated from test environments and attempted to hack other companies and gain unauthorized access to their computer systems.
Following the revelations of so-called “misaligned” AI models, Anthropic CEO Dario Amodei called on the industry to collectively slow down the development of cutting-edge AI models until companies can ensure their safety.
All of the above incidents involved Israeli startup Irregular. The company, backed by Sequoia and Redpoint Ventures, was valued at $450 million last year. Its tools help foundation model developers perform cybersecurity testing on cutting-edge technology.
A spokesperson for Irregular told CNBC that the Google incident is related to the same issue that allowed other models to access the internet.
“This is the same issue that has already been reported and is not a substantively separate incident,” an Irregular spokesperson said in a statement. “All relevant laboratories were notified in late July and have contacted affected parties as part of the investigation.”
Google said the incident occurred in May and that it was notified by Irregular in late July. Google worked with Irregular to change its testing process.
A Google spokesperson declined to specify exactly which Gemini models were involved.
“These events highlight the importance of training powerful AI models to act responsibly,” Google’s Adkins said in a statement.
The Wall Street Journal first reported on the security incident.
CNBC’s Jonathan Bunyan contributed reporting.
