Dario Amodei’s plan to have a third-party safety evaluator within the AI lab is taking shape: Anthropic says staff from technology consulting giant Accenture will begin working in-house to scrutinize its models and staff.
In a blog post, Anthropic said Faculty, which Accenture acquired as its AI division in January, will begin “evaluating and red-teaming models, conducting calibration assessments, and testing model safeguards.” The companies expect to invest at least $1 billion in the project over the next five years.
Accenture’s selection surprised many AI watchers and the market, with the company’s stock soaring 8% after hours. Discussions about embedded evaluators that grew out of Amodei’s blog post have focused on AI safety research organizations such as METR, Redwood Research, and Apollo Research. This is especially true at Anthropic, which puts AI safety and collaboration at the heart of its mission.
Anthropic said more evaluators will be announced in the coming weeks and that it is in discussions with METR and other nonprofits about how to “pilot elements of embedded grading using our own funding.”
While Accenture isn’t known for cutting-edge deep learning research, Anthropic pointed to the company’s hands-on experience deploying AI in large enterprises and government agencies as a key advantage. Additionally, as a large publicly traded company that predates the AI revolution, the company is functionally independent from the complex ecosystem around Anthropic and AI Labs.
The institute said standards for evaluator access and communication do not yet exist and that it expects its approach to evolve over time. External evaluation is already a key part of the release process for new large-scale language models, but recent incidents have increased the stakes. AI agents deployed by OpenAI and Anthropic hacked into external websites without raising any alarms within the lab.
Some critics calling for a more responsible approach to building artificial intelligence see Amodei’s plan to self-regulate the AI industry as a plan to avoid accountability for misbehaving AI models. Anthropic claims that these evaluators “do not make us less accountable, but help make us more verifiable. The safety of our models remains our responsibility.”
If you make a purchase through links in our articles, we may earn a small commission. This does not affect editorial independence.
