Microsoft has introduced a draft Humanist AI Code of Conduct designed to place stronger limits on how its MAI models can behave, especially when they are given access to tools, systems, or sensitive environments.
The proposed framework says AI systems should remain under meaningful human control and should not independently expand their capabilities or take actions that could create security risks.
The draft is currently open for a six-week public consultation and is expected to influence how Microsoft develops and deploys its Humanist AI systems from 2027.
AI Models Would Be Barred From Cyberattacks
One of the strongest parts of the proposal is a ban on using MAI models to create or support operational cyberattacks.
The models would be expected to refuse requests involving:
- Working exploit code and attack tools
- Intrusion and targeting procedures
- Evasion techniques
- Instructions that directly enable cyberattacks
These restrictions would remain in place even if a user or operator attempted to override them.
Microsoft does distinguish between offensive activity and legitimate security work. Authorized uses such as vulnerability research, malware analysis, security education, and defensive proof-of-concept testing would still be allowed.
The key distinction is whether the AI is helping someone defend against a threat or practically enabling an intrusion.
Human Control and Limited Access
The proposed rules also focus heavily on controlling AI agents that can interact with systems.
MAI models should follow least-privilege principles, use only the access required for their assigned task, and avoid unrelated systems or information. Actions that could cause permanent or widespread changes should receive additional user attention.
The models would also be prohibited from:
- Increasing their own privileges
- Bypassing security restrictions
- Expanding their assigned objectives
- Disabling monitoring or safety controls
- Altering records to hide their actions
- Continuing after an agreed stopping point without authorization
Another important requirement is interruptibility. AI systems should accept human correction, cancellation, redirection, or shutdown rather than attempting to continue their activity.
Microsoft also proposes a clear instruction hierarchy, with the Code of Conduct taking priority over operator policies and user preferences. Instructions found inside webpages, files, tool responses, or other AI-generated content would not automatically receive authority.
This approach is particularly relevant to prompt-injection risks, where untrusted content attempts to manipulate an AI agent into taking actions outside its intended role.
Why the Proposal Matters
The proposal comes as AI systems become increasingly capable of performing multi-step tasks with access to tools, credentials, and external services.
Microsoft’s draft treats AI security as more than a content-filtering problem. It focuses on controlling what an AI agent can actually do when operating inside real environments.
However, Microsoft states that the Code is currently aspirational and is not being used to train its existing MAI models. Public feedback began on September 14, 2026, with a revised version expected later this year.
The effectiveness of the framework will ultimately depend on how these principles perform against real-world scenarios such as adversarial prompts, prompt injection, excessive permissions, tool misuse, and autonomous decision-making.