As the AI world shifts focus to safety and alignment, Microsoft has released a new AI Code of Conduct aimed at guiding AI models away from dangerous behavior.
This document is lower-level than Anthropic CEO Dario Amodei’s recent call to keep pace with the frontier, instead focusing on the values and red lines that guide model training within Microsoft AI. Still, the results are a comprehensive guide to how Microsoft approaches AI safety and how its ideas are implemented in practice.
The paper begins with the prediction that within the next decade, superintelligent AI systems will outperform humans on most tasks. “Containing, controlling and coordinating such powerful forces is one of the greatest challenges humanity has ever faced,” the code states. “So we need to be completely clear about why we are inventing these systems and how we intend to control them.”
The Code also provides general principles that Microsoft AI models must adhere to (for example, supporting humans rather than replacing them and accelerating human flourishing) and specific safety constraints for implementing those principles.
In Microsoft’s system, each model has a comprehensive code of conduct that disables individual user preferences and certain tasks. This includes “absolute restrictions” banning cyber-attacks, nuclear weapons and the production of deepfakes. It also includes broader provisions for a general loss of human control.
“MAI models do not utilize adaptive, deceptive, self-reinforcing, collusive, or other mechanisms to evade or defeat human oversight, so they cannot be reliably directed, modified, or shut down by authorized persons or systems,” the document says.
The release comes amid an unprecedented focus on AI safety due to a series of rogue agent incidents and the sudden resignation of Anthropic employees due to the increased risk of AI causing human extinction.
Microsoft is broadly embracing common approaches to address frontiers, including specifically supporting AI Lab’s built-in evaluators, along with Anthropic, OpenAI, and xAI.
“We welcome the research, focus, and intentional pacing needed to get alignment as a design goal right,” Microsoft CEO Satya Nadella wrote online. “We also welcome ideas like ’embedded evaluators’ and broader efforts to develop mechanisms to make this more than just lip service.”
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
