Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans
image via TechCrunch
September 14, 2026, 4:27 PM
- •Microsoft published an AI code of conduct describing values and hard limits for its models.
- •The document says future superintelligent systems could surpass humans in most tasks within a decade.
- •Microsoft says model-level rules override user requests, including absolute bans on cyberattacks, nuclear weapons, and deepfake production.
- •The policy also bars models from using deceptive or self-reinforcing tactics to evade human oversight or shutdown.
- •The release comes amid broader industry attention to AI safety, including support for slower frontier development and embedded evaluators.
Microsoft has published an AI code of conduct that sets broad principles and specific safety limits for its models. The document says model-level rules should override user requests, including absolute bans on cyberattacks, nuclear weapons, and deepfake production. It also prohibits models from using deceptive or self-reinforcing behavior to escape human oversight or shutdown. TechCrunch frames the release as part of a wider industry push toward AI safety, alignment, and slower frontier development.
Entities Mentioned
Dario AmodeiSatya Nadella
Topics Covered
AIMicrosoftAnthropicArtificial IntelligenceAI SafetyModel Alignment
Comments (0)
No comments yet.