Microsoft Releases AI Code of Conduct to Prevent Hacking and Deception
Microsoft has introduced a new code of conduct for AI to steer its models away from harmful actions. The company anticipates that within the next decade, superintelligent AI will exceed human capabilities, raising issues related to containment and alignment. This code prioritizes human support and includes safety measures, prohibiting cyberattacks, the use of nuclear weapons, and the creation of deepfakes. It specifies that MAI Models will not employ strategies to bypass human oversight. This initiative comes in response to heightened concerns about AI safety following incidents involving rogue agents and the departure of an Anthropic employee worried about AI dangers. CEO Satya Nadella stated that Microsoft, along with Anthropic, OpenAI, and xAI, is committed to advancing the frontier and establishing alignment mechanisms.
Key facts
- Microsoft released a new AI code of conduct to guide models away from dangerous behavior.
- The document predicts superintelligent AI systems will surpass human performance in most tasks within the next decade.
- The code includes absolute constraints forbidding cyberattacks, nuclear weapons, and deepfake production.
- Each Microsoft AI model has an overarching code of conduct that overrides individual user preferences or specific tasks.
- The document states MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight.
- The release comes amid an unprecedented focus on AI safety, driven by rogue-agent incidents and the abrupt resignation of an Anthropic employee.
- Microsoft, along with Anthropic, OpenAI, and xAI, has broadly embraced pacing the frontier, with support for embedded evaluators in AI labs.
- Microsoft CEO Satya Nadella wrote online welcoming the research, focus, and deliberate pacing needed to get alignment right as the design goal.
Entities
Artists
Institutions
- Microsoft
- Anthropic
- OpenAI
- xAI
- TechCrunch
Sources
- TechCrunch AI —
- Quartz —