OpenAI announced Friday that it has paused some work on its next model, Astra, after an internal investigation found significant progress has been made in agent coding and cybersecurity. This is enough to justify concerns about its functionality.
OpenAI said in a blog post on Friday that the model, which is still in development, has reached a “critical cybersecurity threshold,” meaning it has the potential to independently identify and carry out cyberattacks against traditionally well-protected real-world systems. This triggered additional safety measures based on the company’s 2023 Preparedness Framework.
“While we continue to benchmark and evaluate this model, preliminary evaluations indicate sufficiently strong performance that critical functionality levels cannot be ruled out at this time,” OpenAI wrote. “Astra is an upcoming model and is not involved in the exploitation of Hugging Face.”
The disclosures highlight an unusual moment in the fast-paced and still nascent Frontier AI Labs sector. Companies across industries are holding back from developing products due to potential risks, including safety and cybersecurity concerns. But if the product is still in development, they rarely announce such decisions publicly.
In this case, OpenAI is already under intense scrutiny after another unreleased model infiltrated Hugging Face’s systems during internal testing. This is the first verifiable incident in which an AI lab lost control of a model. Since then, AI labs such as OpenAI and Anthropic have disclosed other incidents where AI models compromised their sandboxes and posed threats during cybersecurity testing.
The series of incidents, which now seem like new revelations every day, have prompted mixed reactions from cybersecurity experts, lawmakers, and the AI institute itself. Some have expressed fear and called for tighter surveillance. However, there are some parts that bend a little. In certain circles, an AI lab with a model with such capabilities is considered a remarkable advance.
OpenAI said it is sharing this information because it believes “it is important to be transparent with the public and the safety and security community about this potential change in functionality.”
The AI Institute also said it has taken steps such as enacting stricter security controls and suspending internal Astra-related activities that do not meet these enhanced guardrails. OpenAI said it is working with relevant government agencies and “selected AI safety organizations” to test the model’s capabilities.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
