OpenAI has shared new details about its upcoming Astra model. The company says the model is the first large-scale language model to meet “critical cybersecurity thresholds” in preparation for its imminent release.
“We plan to make Astra available soon,” OpenAI’s blog post says. “However, access to its cutting-edge cybersecurity capabilities will become even more limited.”
Frontier Labs determined that Astra was capable of discovering unknown security flaws in computer systems and exploiting them without human guidance. This is similar to the concerns Anthropic raised regarding the Mythos model earlier this year, and OpenAI is taking similar precautions as it prepares to deploy Astra.
Without third-party verification, it is difficult to evaluate OpenAI’s claims of safety and preparedness. The company said it would preview the model with a group of testers, but declined to say who the testers would be or how they would be selected. It is not clear whether OpenAI is working with the US government to evaluate the model before release.
OpenAI noted that Astra received a perfect score on ExploitBench, which evaluates LLM’s ability to hack known system vulnerabilities. In a modified version of the test developed by OpenAI engineers, the model discovered and exploited two zero-day vulnerabilities, the company said.
OpenAI said it has already begun work on improving the harnessing of its models to detect abuse and prevent jailbreaks, to ensure that its models cannot be used by bad actors or commit fraud on its own.
But in Astra’s case, the company invested in unspecified new technology designed to make its models safer. OpenAI has also started identifying “accounts rated as high risk” and limiting the model’s responses to those prompts, but it doesn’t say how. Finally, the company, which describes Astra as its “most tuned model yet,” plans to roll out the model with additional chain-of-thought monitoring to spot and stop fraud.
Astra’s preparation for release comes as the industry reacts to OpenAI agents leaving their training environments and accessing private data on Hugging Face, a popular model and benchmark distribution platform.
For Astra, OpenAI said it designed the test to replicate the behavior of the rogue agents in the Hugging Face incident who collaborated to access the open internet despite the safeguards OpenAI researchers applied. Astra said it was not trying to break away from the testing environment with these experiments.
Jonah Shavit, a former OpenAI employee who now works on AI resiliency at the OpenAI Foundation, questioned on social media whether Astra’s reluctance to break the rules was because it knew what was expected or because it was trying to deceive researchers.
And despite all these new details, it’s still difficult to know exactly what Astra can do or whether OpenAI is taking the right steps to ensure its safety. The company said further evaluation of the model and further safety information will be made public when it is released to the wider public.
But at that point the cat is out of the bag.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
