OpenAI on Tuesday announced a set of security policies focused on containing security incidents during model testing. New safety measures include closer monitoring of models during the development process, as well as increased tuning and security during the post-training process.
“As the capabilities of our models increase, so too does the risk associated with in-house development and testing,” the company said in a blog post. “Our standards for monitoring, collaboration, and security must stay ahead of these risks.”
This new measure is one of the first public changes in OpenAI’s safety practices since shortly after the face-hugging incident that came to light on July 21st.
OpenAI representatives said the action was not a direct response to the Hugging Face incident, but was also prompted by the cybersecurity features of the next Astra model and the overall pace of advancement in AI development.
In the same post, OpenAI revealed that it paused reinforcement learning (RL) for two weeks following the face-hugging incident, but has since restarted many of its low-risk models.
“While our largest planned frontier RL run remains on hold, we will conduct small-scale training and evaluation to evaluate model behavior, validate safeguards, and establish more evidence of consistency before proceeding,” the post reads.
Amelia Grace, OpenAI’s vice president of research, told reporters that as models grow in power, controls become more stringent, with the largest models facing the most scrutiny.
“We have set requirements and expectations for safe development,” Grace told reporters. “These requirements and expectations vary depending on the level of risk we perceive.”
OpenAI has been criticized for insufficient network security efforts after a model was able to escape from its training environment by compromising a tool on an Internet-accessible network. New safety measures include stronger network isolation practices, but details remain vague. Under the new system, “a single compromise of a workload or support service does not grant unauthorized access to the Internet or other internal networks,” the post said.
The most powerful safeguard is a monitoring system that examines tool behavior, available inference traces, and activity logs for various fraudulent activities. OpenAI says it aims to issue alerts within 30 minutes of concerning activity occurring.
OpenAI estimates that the computing load of its monitoring is approximately 20% of the monitored process. The company promised more details about the system in a future blog post. The company’s official post-mortem analysis of the incident is also still pending.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
