Microsoft's new AI 'code of conduct' tells models not to hack systems or trick humans
Microsoft published a formal AI code of conduct outlining values and hard constraints for its AI models, including absolute prohibitions on cyberattacks, nuclear weapons assistance, and deepfake production. The document predicts superint…
- 01The document predicts superintelligent AI surpassing human performance across most tasks within a decade and frames alignment as one of humanity's greatest challenges.
- 02The code establishes a hierarchy where model-level conduct rules override individual user preferences, explicitly barring models from using deceptive or self-reinforcing mechanisms to evade human oversight or shutdown.
- 03The release follows a series of rogue-agent incidents and an Anthropic researcher's resignation over extinction risk concerns.
- 04Microsoft CEO Satya Nadella publicly endorsed the broader industry shift toward deliberate pacing and embedded evaluators as accountability mechanisms.
Microsoft published a formal AI code of conduct outlining values and hard constraints for its AI models, including absolute prohibitions on cyberattacks, nuclear weapons assistance, and deepfake production. The document predicts superintelligent AI surpassing human performance across most tasks within a decade and frames alignment as one of humanity's greatest challenges. The code establishes a hierarchy where model-level conduct rules override individual user preferences, explicitly barring models from using deceptive or self-reinforcing mechanisms to evade human oversight or shutdown.
Read the full article at techcrunch.comMicrosoft published a formal AI code of conduct outlining values and hard constraints for its AI models, including absolute prohibitions on cyberattacks, nuclear weapons assistance, and deepfake production. The document predicts superintelligent AI surpassing human performance across most tasks within a decade and frames alignment as one of humanity's greatest challenges. The code establishes a hierarchy where model-level conduct rules override individual user preferences, explicitly barring models from using deceptive or self-reinforcing mechanisms to evade human oversight or shutdown. The release follows a series of rogue-agent incidents and an Anthropic researcher's resignation over extinction risk concerns. Microsoft CEO Satya Nadella publicly endorsed the broader industry shift toward deliberate pacing and embedded evaluators as accountability mechanisms.
Don't miss tomorrow's
The Daily Pulse in your inbox each morning — sourced and linked.