OpenAI’s new Astra model uses a reasoning technique called “recurrent depth” that allows it to operate outside of the sequential thinking that characterizes most reasoning models, The Information reported on Tuesday. The technique, also known as “opaque recurrence,” is likely to make monitoring the model’s chain of thought even more difficult, and has AI safety experts perturbed.
While Astra’s use of the technology is reportedly limited, its emergence still raises major concerns among AI safety experts.
“We are extremely concerned about reports of Astra’s opaque relapse,” Redwood CEO Buck Schlegelis said in a post after the news broke. “We don’t know if Astra will make CoT observable much less than previous models. However, if OpenAI pushes this technology further, it will give us the option to significantly increase recurrence and completely destroy CoT observability.”
Zvi Mowshowitz, a longtime AI safety advocate, echoed similar sentiments, writing that legislation may be needed to prevent a “race to the bottom” among AI labs.
“This technology is playing with fire and jeopardizes the taboos that OpenAI and Anthropic have fought to establish and work hard to maintain chain of thought fidelity and observability for as long as possible,” Mowshowitz wrote. “More intensive use of such technology will likely impair observability.”
Under normal circumstances, a reasoning model’s chain of thought provides a sequence of steps for the model to take as it attempts to solve a problem. Although this representation is incomplete, it serves as a valuable tool for monitoring fraud and misalignment. In the case of OpenAI’s recent rogue agent activity, chain-of-thought records were an important tool in uncovering why the agent acted the way it did.
In opaque iteration, the model takes a nonlinear approach and processes the same query multiple times in a loop. The result is fewer legible traces, effectively avoiding traditional chain-of-thought records.
Importantly, Astra’s use of this technology appears to be limited. The model’s chain of thought is expected to remain legible, and the company has balked at any suggestion of moving to “neuralize.” OpenAI has already announced plans for an extensive chain of thought monitoring system as part of its forward-looking safety plans.
In a post on X, Jakub Pachocki, chief scientist at OpenAI, highlighted the institute’s commitment to easy-to-read thought chains. “OpenAI has been committed to preserving and leveraging chain-of-thought monitoring since our first inference models,” Paciocchi wrote. “This is a core goal of our current research program.
All AI models perform some degree of opaque reasoning, but few researchers view thought chain logs as a direct representation of the model’s reasoning. Still, these caveats do not allay concerns that opaque recurrences could make monitoring AI inference difficult, especially as usage increases across different models. In a follow-up report Wednesday morning, The Information reported that both Anthropic and Google DeepMind are already in talks about the technology.
Ryan Greenblatt, chief scientist at the Redwood Institute, said in a post reacting to the news that opaque reasoning is more easily extended than traditional chain-of-thought reasoning and could effectively eliminate all reasoning from visible channels.
“My biggest concern is that the natural progression from here involves scaling up opaque inference to the point where the model infers entirely, or almost entirely, in latent space,” Greenblatt wrote. “We hope it’s not too late to avoid the most concerning architectures and that OpenAI stops here.”
If you make a purchase through links in our articles, we may earn a small commission. This does not affect editorial independence.
