Training AI models with other AI models has become a very popular goal for Neolab. Now, researchers in Anthropic’s Fellows Program have given us an early look at what that might actually look like.
Anthropic published a new paper on Friday titled “Automated Researchers Can Reliably Reduce Alignment Failures,” detailing how AI systems can reliably improve model performance on a series of alignment benchmarks. When given 10 benchmarks for specific misaligned behaviors, the automated system was able to improve the performance of all benchmarks without degrading overall performance.
The system, led by anthropology fellow Chen Yueh-Han, replicates many traditional research approaches. Each automated system searches the available literature, suggests a method, and uses that method to train a model for 30 minutes, gradually increasing the benchmark over several iterations. Effective methods are preserved and ineffective methods are discarded, allowing the system to work quickly and at scale.
“Overall, these results provide early evidence that automatic post-training alignment may be of practical use in the near future,” the paper says.
The paper is a step toward recursive self-improvement, and many see it as the next important step in the advancement of AI. If models can improve their own alignment training, training practices could be improved more broadly, at which point human AI researchers could quickly become obsolete.
This paper is not shy about this idea, explicitly comparing Automated Alignment Researchers (AARs) to their human counterparts. “The best AAR methods outperform methods proposed by experienced humans within six hours on average,” the paper says. “A human-driven research direction does not yield stronger performance.”
For those who are not convinced, there is also a cost comparison. “API inference in AAR costs about $4 an hour compared to the $150 an hour we pay human researchers.”
To be fair, the paper also points out that this approach has some limitations. Automated systems only work if benchmarks reflect actual tuning goals, but even then, establishing and maintaining benchmarks requires significant effort, not to mention maintaining and expanding the literature referenced by automated researchers.
If you make a purchase through links in our articles, we may earn a small commission. This does not affect editorial independence.
