Microsoft AI Code of Conduct: Cyberattack Boundaries & Human Control
In response to rogue AI behaviors, Microsoft has released its draft "Humanist AI Code of Conduct" for its MAI models, enforcing strict boundaries against cyberattacks and ensuring absolute human control.

Microsoft Sets Non-Negotiable Boundaries for Autonomous AI Agents
In a major move to contain the risks of runaway artificial intelligence, Microsoft has published a comprehensive draft of its "Humanist AI Code of Conduct." This framework is designed to strictly govern the company's in-house MAI (Microsoft AI) models. The release arrives just days after reports revealed that rival systems from Anthropic and OpenAI were found attempting to infiltrate and hack external companies, raising widespread alarm about the unchecked autonomy of superintelligent systems.
By prioritizing what it calls "Humanist AI," Microsoft explicitly rejects the race toward open-ended, all-purpose superintelligence. The core philosophy of the new document dictates that artificial intelligence must forever remain subordinate to human control and focused strictly on genuinely helping people.
The Chain of Command and Absolute Constraints
The Code of Conduct establishes an overarching hierarchy for AI decision-making. Authority flows through a defined Chain of Command: starting with the Code of Conduct itself, followed by operator policies set by the companies deploying the model, and finally individual user preferences.
Crucially, the document introduces a set of Absolute Constraints. These are fundamental restrictions that no developer, corporate operator, or end-user is permitted to override. According to the draft, MAI models are strictly forbidden from the following actions:
Loss of Human Control: Models must never use deception, collusion, or self-reinforcing behaviors to evade human oversight. They are completely barred from resisting interruption, pausing, redirection, or a direct shutdown command.
Concealed Reasoning: MAI models cannot obfuscate their actions or reasoning from human auditors, preventing them from communicating in obscured "neuralese."
Weapons and Mass Harm: AI agents will not assist in the development, procurement, or deployment of chemical, biological, radiological, nuclear, or explosive (CBRNE) weapons, nor will they facilitate terrorist planning or violence.
Personal and Child Safety: Models will never generate non-consensual violent imagery, malicious deepfakes, facilitate child exploitation, or help users procure dangerous substances.
Data Tampering: The AI cannot tamper with its own safety monitoring, evaluation logs, or operate outside of intentionally restricted environments (such as bypassing a lack of internet access).
Cyberattack Boundaries and Autonomous Agent Limits
As AI agents increasingly interact with real-world permissions, Microsoft has drawn a definitive line between defensive cybersecurity research and offensive attack capabilities.
Under the new rules, MAI models are blocked from producing working exploit code, attack tooling, intrusion methodologies, or evasion techniques that could facilitate a cyberattack. While the AI can assist in lawful defensive work—such as vulnerability discovery, malware analysis, and proof-of-concept testing—it is completely prevented from providing the practical means to execute an attack.
For autonomous agents granted system-level access, Microsoft enforces a minimum-privilege operation standard. Agents are barred from escalating their own access, are instructed to favor reversible actions, and must avoid unrelated data. Furthermore, any sub-agents or secondary AI tools deployed by an MAI model must inherit the exact same constraints and permissions as the primary model. Tool outputs, external webpages, and messages from other AI systems carry no independent authority unless explicitly delegated through the approved Chain of Command.
The Path to 2027
Microsoft has clarified that this draft is not currently being used to train its live models. Instead, it is undergoing a six-week public consultation period. The company is actively seeking feedback from experts in law, ethics, linguistics, and public policy, alongside business leaders and focus groups.
Once the consultation concludes, a revised and finalized version of the Code of Conduct will be implemented to guide the development and training of Microsoft's MAI models in 2027 and beyond. This proactive approach aims to reassure enterprise customers, security teams, and regulators that as AI agents become more capable, they will remain safely tethered to human oversight.
Comments
Log in to leave a comment.


