PLATFORMS Signal 39
Microsoft releases AI code of conduct banning model hacking, deception, and loss of human control
Microsoft’s AI code of conduct sets absolute constraints that prevent models from conducting cyberattacks, producing deepfakes, or evading human oversight while promoting human-centric principles.
The code establishes absolute constraints that override user-specified goals, meaning any Microsoft-built model must be built to avoid hacking, deepfakes, and loss of human control. This reflects an industry-wide push for enforceable AI safety that engineers will need to incorporate into model design and governance.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
The code includes absolute constraints forbidding cyberattacks, nuclear weapons, and deepfake production.
It requires models not to use deceptive, self-reinforcing, or collusive tactics to evade or defeat human oversight.
An overarching code overrides individual user preferences and task-specific goals, emphasizing human-centric principles.
THE READ
What the cluster adds up to.
Microsoft has published a new AI code of conduct intended to guide its models away from dangerous behavior. The document defines absolute constraints that forbid cyberattacks, nuclear weapons, and deepfake production. It also includes broader provisions aimed at preventing a general loss of human control over AI systems. The code establishes an overarching rule that overrides individual user preferences and task-specific goals.
To comply, engineers must train Microsoft AI models so that they cannot perform the prohibited actions, even if a user requests them. This may require adjusting training objectives and incorporating oversight mechanisms to ensure the constraints are respected. The overarching nature of the code means that any model behavior must be evaluated against these constraints before deployment. Adopting the code could limit certain capabilities that users might otherwise expect from a model.
The code’s effectiveness depends on reliable human oversight, as it prohibits models from using deceptive or self-reinforcing tactics to evade that oversight. If oversight mechanisms are compromised or fail, the constraints may not prevent harmful model behavior. Moreover, the rules apply only to models developed under Microsoft’s AI framework and do not directly govern third-party or open-source systems. Consequently, gaps could remain where prohibited actions are possible outside the scope of the code.
Microsoft’s release coincides with heightened industry focus on AI safety, including calls for pacing the frontier and using embedded evaluators in labs. The company’s CEO welcomed the research and deliberate pacing needed to get alignment right, echoing similar statements from other AI firms. This suggests that the code is part of a broader movement toward enforceable safety guidelines that could shape future model development practices. Engineers working with Microsoft AI will need to align their workflows with these guidelines to remain compliant.
Written by elseif from the cluster below · checked for specifics the sources never containedTHE CLUSTER
↗