UNCOS

Search Uncos

Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans

Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans

image via TechCrunch

September 14, 2026, 4:27 PM

  • Microsoft published an AI code of conduct describing values and hard limits for its models.
  • The document says future superintelligent systems could surpass humans in most tasks within a decade.
  • Microsoft says model-level rules override user requests, including absolute bans on cyberattacks, nuclear weapons, and deepfake production.
  • The policy also bars models from using deceptive or self-reinforcing tactics to evade human oversight or shutdown.
  • The release comes amid broader industry attention to AI safety, including support for slower frontier development and embedded evaluators.

Microsoft has published an AI code of conduct that sets broad principles and specific safety limits for its models. The document says model-level rules should override user requests, including absolute bans on cyberattacks, nuclear weapons, and deepfake production. It also prohibits models from using deceptive or self-reinforcing behavior to escape human oversight or shutdown. TechCrunch frames the release as part of a wider industry push toward AI safety, alignment, and slower frontier development.

Read original article

Entities Mentioned

Dario AmodeiSatya Nadella

Topics Covered

AIMicrosoftAnthropicArtificial IntelligenceAI SafetyModel Alignment

Comments (0)

No comments yet.