Microsoft CEO Satya Nadella is the latest tech executive to give a long think about how AI can be made safer.
In a post on X Saturday morning, Nadella wrote that it was time to “step back and evaluate the trust architecture” of AI.
“We cannot treat superintelligence as a set of nested black boxes whose recommendations, answers, and actions cannot be simply accepted or rejected,” Nadella wrote, using the Trump administration’s preferred AI terminology.
As Nadella outlined, this approach “means separating the model from the harness that coordinates its work” and also “externalizes controls and safeguards.” He also called for a system that would document “all meaningful model actions” with “tamper-proof human-readable evidence,” and that “authorized individuals” have the ability to “pause or shut down the model mid-task” at any time.
“We need to assume the model is compromised and contain it from the beginning,” he said. “Think of it like an emergency brake.”
Nadella’s comments came as major AI companies acknowledged a growing number of incidents in which they appeared to have lost control of their models, and Anthropic CEO Dario Amodei announced more cautious AI development plans.
