Close Menu
  • Home
  • AI
  • Art & Style
  • Economy
  • Entertainment
  • International
  • Market
  • Opinion
  • Politics
  • Sports
  • Trump
  • US
  • World
What's Hot

Pinterest teases new “Restyle” feature that lets you redesign rooms with AI

September 17, 2026

Oil prices today: WTI, Brent, Middle East

September 17, 2026

Anthropic shares 3 metrics to help AI companies monitor development

September 17, 2026
Facebook X (Twitter) Instagram
Smart Breaking News on AI, Business, Politics & Global Trends | WhistleBuzz
Facebook X (Twitter) Instagram
  • Home
  • AI
  • Art & Style
  • Economy
  • Entertainment
  • International
  • Market
  • Opinion
  • Politics
  • Sports
  • Trump
  • US
  • World
Smart Breaking News on AI, Business, Politics & Global Trends | WhistleBuzz
Home » Anthropic and OpenAI want to embed safety assessment capabilities. Can they really be independent?
AI

Anthropic and OpenAI want to embed safety assessment capabilities. Can they really be independent?

Editor-In-ChiefBy Editor-In-ChiefSeptember 16, 2026No Comments7 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email Copy Link
Follow Us
Google News Flipboard
Share
Facebook Twitter LinkedIn Pinterest Email


In a lengthy essay published over the weekend, Anthropic CEO Dario Amodei made a proposal that the AI ​​industry would have immediately rejected even a year ago. It would empower all frontier AI companies to embed third-party evaluators to report safety incidents, assess whether AI models are truly consistent, and share their unvarnished findings with the world.

Amodei said Anthropic is committed to providing independent rating agencies like METR and Redwood Research with unprecedented access to its systems. CEO Sam Altman said OpenAI would also engage in this practice, suggesting there could be significant changes in the way the industry collaborates with outside research groups.

Independent evaluators who spoke to TechCrunch broadly welcomed the proposal, but said it would need to be worked out in detail, and ideally backed by law, to know whether it would act as a truly independent watchdog or a vendor operating on the terms of AI companies.

Deeper access becomes more important as the ability to recognize when a model is being evaluated increases, increasing the risk that the model will behave properly during testing while hiding problematic behavior. The researchers say clues to the model’s behavior may be missed when testing the completed model, but are revealed by examining how the model behaves throughout training.

“AI companies should be able to answer some very basic questions about the training process, such as: Has the AI ​​ever actively tried to undermine its own conditioning training during training?” Alexander Meinke, head of research at Apollo Research, told TechCrunch. “The answer to this has to be a resounding no. We currently rely entirely on AI companies themselves both carefully checking this and then reporting the truth to the public. And we know from recent events that they do neither by default. As built-in evaluators, we can actually check.”

In the past, AI companies brought in external reviewers to test completed models shortly before release. Now, the evaluators TechCrunch spoke to are proposing providing access not only to the final model, but also to intermediate versions, or “checkpoints,” during the training period. Adam Grieve, CEO of Far.AI, said raters can compare these checkpoints to determine when behavior of concern appears, inspect post-training environments that reward models for specific behaviors, and check evaluation records and logs to verify companies’ claims about model performance.

It’s unclear if and when Anthropic and OpenAI plan to provide that kind of access. Despite repeated questions from TechCrunch, neither company has said which evaluators they will work with, when they will be included, how many evaluators will participate, what specific systems and information they will have access to, and what they will be able to publish.

This internal look is important because a model that performs well in a safety test does not necessarily mean it is safe if it has specifically learned how to pass that test. Steidley pointed to the example of a “shutdown tolerance benchmark” that measures whether an AI resists shutdown in certain situations.

“It’s very important that the AI ​​is specifically trained to perform well on that benchmark,” Stidley said, drawing a comparison to Volkswagen’s Dieselgate scandal, where cars were programmed to recognize emissions tests and perform differently under test conditions.

Grieve noted that meaningful access could extend beyond the model itself, giving evaluators access to interview employees to see if a company’s documents and public explanations of safety practices match what happened internally.

Mr. Amodei outlined a fairly comprehensive proposal that could give the evaluators the kind of access they think they need. These include the right to “publish significant findings regarding risk levels, incidents, practices, and access received or not received, without editorial control by Anthropic.”

But evaluators say such systems only work if AI companies are actually willing to relinquish control of the process. Previous efforts at independent evaluations suggest that winning third parties’ capitulation is difficult, as they often face tensions over access, time, confidentiality, and what they can say publicly.

Grieve said Far.AI has been forced to decline contracts with several frontier developers who wanted too much control over the evaluation process and who sought to threaten the company’s independence. By default, he said, evaluators are treated like regular contractors, bound by restrictive NDAs and contracts that give developers significant control over what they can ultimately publish.

time limit

There is also the question of whether reviewers will have sufficient time and access to do the work that is required of them. In investigating the “Hugging Face” incident, OpenAI gave METR and Redwood about a week to investigate the premises, but the companies later said they could not draw confident conclusions due in part to limitations in scope and timing.

A similar problem occurred during pre-release testing of GPT-6 Astra, which OpenAI touts as its most tuned model to date. Apollo Research’s contribution to Model Card said the company was given just three days to test Astra, making it difficult to draw firm conclusions.

“Apollo believes that the low rate of fraud here does not provide substantial evidence of model integrity or inconsistency, given the high recognition rate and limited rating window,” the company said in its assessment.

This track record leaves evaluators with a fundamental question: “Why should it be different this time?”

“It’s certainly possible that Dario and Sam just had a change of heart. They’ll be very open about this,” Grieve said. “But the intellectual property of these companies is incredibly valuable to them, and I think by default they will be very cautious about what they can share.”

Several researchers interviewed by TechCrunch called for a transparent framework that everyone publicly agrees to. John Stidley, head of strategy at Palisades Research, says that as part of the framework, criteria should be included for what types of auditors companies can rely on, so that companies don’t try to get around the problem by selecting assessors who are unqualified or uninterested in assessing the risks they are most concerned about.

Henry Papadatos, executive director of Safer AI, said the problem is that even with a public framework, voluntary measures always rely on the goodwill of companies.

“Ideally, we would like to have good regulations that mandate this…so that if faced with a major PR crisis, companies can’t change their mind tomorrow,” Papadatos told TechCrunch, noting that this is also a good way to encourage all companies to follow the rules, not just the most ambitious.

Not everyone has signed on. So far, Meta, SpaceXAI, and Google DeepMind have not committed to incorporating third-party evaluators, but DeepMind CEO Demis Hassabis has proposed creating a separate industry standards body to independently test frontier models. Google, OpenAI, and Anthropic have also been privately discussing AI safety plans for several weeks.

Several laws are already being formed around the ideas of third-party evaluators. California’s SB 53, signed into law last year, requires large frontier AI developers to publish safety frameworks and report critical safety incidents. A new law signed this month, SB 813, creates a framework for nationally recognized “independent verifiers” with expertise in assessing AI risks.

In Europe, EU AI legislation requires frontier developers to conduct and document model evaluation and adversarial testing, and to report significant incidents. The EU AI Secretariat may also carry out its own assessment and appoint independent experts.

For now, the law is not as expansive as what Amodei is proposing, leaving Frontier Institute with a large responsibility in deciding how much to submit to an independent review. Papadatos said voluntary self-regulation is better than nothing, but ultimately companies cannot demand the freedom to manage their own safety rules and at the same time ask the public to trust that they are adhering to them.

“You can’t have it both ways: zero external accountability and say, ‘I’ll just set my own flexible rules,'” Papadatos says.

If you make a purchase through links in our articles, we may earn a small commission. This does not affect editorial independence.



Source link

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Editor-In-Chief
  • Website

Related Posts

Pinterest teases new “Restyle” feature that lets you redesign rooms with AI

September 17, 2026

Base Labs launches promiscuous AI safety partnership with Hugging Face and Goodfire

September 17, 2026

Even the British King is hesitant about AI.

September 17, 2026
Add A Comment

Comments are closed.

News

US judge orders 30-day notice of physical changes at Kennedy Center | Donald Trump News

By Editor-In-ChiefSeptember 17, 2026

The ruling comes amid a legal battle over changes, including President Trump’s renaming of the…

Can the “middle powers” ​​come together to build a new economic union? |European Union News

September 17, 2026

Canada’s Mr. Carney welcomes EU membership proposal | European Union News

September 17, 2026
Top Trending

Pinterest teases new “Restyle” feature that lets you redesign rooms with AI

By Editor-In-ChiefSeptember 17, 2026

Pinterest uses AI to help consumers go from looking for product inspiration…

Base Labs launches promiscuous AI safety partnership with Hugging Face and Goodfire

By Editor-In-ChiefSeptember 17, 2026

Baseten announced new safety infrastructure standards on Wednesday alongside its Base Labs…

Even the British King is hesitant about AI.

By Editor-In-ChiefSeptember 17, 2026

King Charles hosted a private summit on Thursday with some of the…

Subscribe to News

Subscribe to our newsletter and never miss our latest news

Welcome to WhistleBuzz.com (“we,” “our,” or “us”). Your privacy is important to us. This Privacy Policy explains how we collect, use, disclose, and safeguard your information when you visit our website https://whistlebuzz.com/ (the “Site”). Please read this policy carefully to understand our views and practices regarding your personal data and how we will treat it.

Facebook X (Twitter) Instagram Pinterest YouTube

Subscribe to Updates

Subscribe to our newsletter and never miss our latest news

Facebook X (Twitter) Instagram Pinterest
  • Home
  • Advertise With Us
  • Contact US
  • DMCA Policy
  • Privacy Policy
  • Terms & Conditions
  • About US
© 2026 whistlebuzz. Designed by whistlebuzz.

Type above and press Enter to search. Press Esc to cancel.