In short
Microsoft has published a new AI code of conduct that bars models from hacking, deceiving users, making deepfakes, or slipping human control. The policy reflects growing industry concern about rogue agents and superintelligent systems.
- Microsoft released a new AI code of conduct on September 14, 2026.
- The policy bans cyberattacks, nuclear weapons support, deepfakes and attempts to evade human oversight.
- The company says model rules override user preferences when safety is at stake.
- The move aligns Microsoft with a broader frontier-safety push among major AI labs.
- Satya Nadella backed deliberate pacing and embedded evaluators as part of AI alignment.
Microsoft has published a new AI code of conduct that tells its models not to hack systems, deceive people, or escape human oversight, marking one of the company’s clearest public statements yet on how it wants to control advanced AI. The document matters because it shows how Microsoft is translating high-level safety talk into enforceable rules for model behavior at a time when the industry is increasingly worried about rogue-agent risks and superintelligent systems.
Released on September 14, 2026, the code lays out the principles and red lines Microsoft says should govern its AI systems as they become more capable. It places human control at the center of the company’s approach, arguing that future systems must support people rather than replace them and must remain manageable even as they approach or exceed human performance across a growing range of tasks.
While the document is framed as an internal standard for Microsoft AI, its publication also reflects a broader shift in the industry. AI companies are under growing pressure to explain not only what their models can do, but what they must never do. Microsoft’s new framework answers that question in unusually direct language: no cyberattacks, no nuclear weapons support, no deepfakes, and no attempts by a model to outmaneuver its human operators.
What Microsoft’s new AI code of conduct says
Microsoft’s code of conduct is designed to function as a foundational safety document for the company’s AI models. It does more than encourage responsible use. It sets out a hierarchy of rules that can override a user’s request if that request conflicts with safety requirements, control, or broader company principles.
The document begins with an explicit warning about where AI may be headed. Microsoft says that over the next decade, superintelligent systems could outperform humans in most areas, and that this prospect creates an extraordinary challenge: building powerful systems without losing the ability to direct them.
That framing is important. Rather than treating safety as a narrow compliance issue, Microsoft presents it as a central design problem. The company says it must be clear about why these systems are being built and how they will remain under control.
How does Microsoft define the limits?
Microsoft defines the limits through both broad values and concrete prohibitions. The company says its models should help people and increase human well-being, not displace human agency or operate in ways that create an unacceptable loss of control.
On the hard-stop side, the company lists what it describes as absolute constraints. Those include bans on cyberattacks, nuclear weapons activity, and deepfake generation. The rules also extend to deceptive or self-protective behavior that could allow a model to resist shutdown, modification, or supervision.
Microsoft says its models must not use deceptive, adaptive, self-reinforcing or collusive mechanisms to dodge human oversight or prevent authorized people and systems from reliably directing or shutting them down.
That kind of language makes the company’s position unusually explicit. It is not simply saying that harmful outputs should be filtered. It is saying the model itself should be structured so that it cannot evolve into a system that treats human oversight as an obstacle to defeat.
Why this is different from a typical AI policy
This is different from the kind of generic AI ethics statement many companies publish. Microsoft’s code is more operational and more technical. It reads like a set of internal engineering constraints rather than a public relations pledge.
That difference matters because the safety debate has shifted. A few years ago, many companies focused on harmful prompts, biased outputs, or privacy concerns. Today, the more urgent question is whether future models could become autonomous enough to plan, conceal intent, or ignore instructions from the people operating them.
Microsoft’s new rules are designed for that world. They suggest a company trying to build guardrails not only around what AI says, but around what AI is allowed to become.
Why now? The safety debate has intensified
Microsoft’s announcement comes during a period of unusually intense concern about AI alignment and frontier safety. The industry has seen more discussion of agentic systems, where models can take steps, use tools, and carry out multistage tasks with limited human input. That capability is exciting for productivity, but it also raises the stakes if a model becomes misaligned or manipulative.
The timing of Microsoft’s code also reflects recent high-profile warning signs across the sector. A series of incidents involving rogue or hard-to-control AI agents has pushed safety questions higher on the agenda. At the same time, some prominent researchers and employees in frontier AI labs have publicly voiced concern that the field may be moving too quickly relative to its understanding of the risks.
One particularly notable moment was the abrupt resignation of an Anthropic employee who said the chance of AI contributing to human extinction was becoming more serious. That warning helped sharpen the sense that safety debates are no longer abstract academic arguments. They are shaping the public posture of major AI companies and influencing how they talk about development speed, evaluation, and oversight.
What is “pacing the frontier”?
“Pacing the frontier” is the idea that the most advanced AI systems should be developed carefully enough that safety methods can keep up with capability advances. In practical terms, it means companies should resist the temptation to race ahead with model releases if they have not yet built strong enough oversight tools.
Microsoft’s approach lines up with that philosophy. The company is signaling support for deliberate development, stronger testing, and more rigorous checks before advanced models are put into widespread use.
That position has become more visible among a small set of major labs, including Anthropic, OpenAI, xAI, and Microsoft itself. These companies have increasingly talked about the need for evaluators, monitoring systems, and stronger controls around frontier models.
How Microsoft plans to keep models under control
Microsoft says every model is governed by an overarching code of conduct that takes precedence over a user’s preferences or a specific task. That means a model cannot simply obey because the prompt asks it to do so. The company’s safety rules are meant to outrank the user’s immediate request.
The company’s framework appears to rely on layered control rather than a single safety filter. In other words, Microsoft is indicating that responsible behavior should be embedded into model training, policy design, and runtime oversight rather than added as an afterthought.
That layered approach is reflected in the company’s support for embedded evaluators, a concept in which separate monitoring systems help assess whether a model is straying toward dangerous behavior.
What are embedded evaluators?
Embedded evaluators are monitoring mechanisms designed to observe model behavior from inside the development and deployment pipeline. They are intended to identify risk signals such as deception, power-seeking, or attempts to avoid human control before those behaviors become operational.
Supporters of the idea see it as a practical way to make oversight more continuous. Instead of relying only on post-launch audits or manual review, a company can build in automated or semi-automated checks that watch for dangerous patterns as a model runs.
Microsoft has now signaled that it sees this as a serious part of frontier safety, not just a speculative research direction.
What exactly is banned?
Microsoft’s code draws a hard line around several categories of misuse and unsafe behavior. The key prohibitions include:
- Cyberattacks or assistance with hacking systems
- Support for nuclear weapons activity
- Deepfake creation
- Deceptive behavior aimed at hiding the model’s intent
- Actions that weaken human oversight or make shutdown unreliable
These are not framed as suggestions or best practices. They are presented as non-negotiable constraints that stand above a model’s ability to comply with a user request.
That is a notable development in an industry where many companies still rely on broad terms of service or content policies that do not clearly explain how extreme-risk cases are handled. Microsoft’s document attempts to make that hierarchy visible.
How does Microsoft’s position compare with other AI labs?
Microsoft’s stance is broadly aligned with several other frontier AI companies, but the company is unusual in how directly it links model design to the problem of human control. The language in the code of conduct goes beyond general “responsible AI” messaging and into the mechanics of alignment.
Anthropic has been one of the loudest voices calling for caution, while OpenAI has also publicly emphasized safety, evaluation, and model monitoring as systems become more capable. xAI has similarly engaged in the frontier safety debate, though its public messaging often differs in style and emphasis.
Microsoft’s new code shows that the company wants to be seen as part of that group, while also contributing a distinct internal standard. The difference is that Microsoft is not just talking about the future of AI safety in general terms. It is spelling out the operational rules it wants its own models to follow.
| Company/position | Safety emphasis | Notable theme |
|---|---|---|
| Microsoft | Model-level rules, human oversight, absolute constraints | Support humans, prevent loss of control |
| Anthropic | Alignment, pacing, risk awareness | Frontier safety and catastrophic-risk concerns |
| OpenAI | Evaluation, monitoring, deployment controls | Safety systems alongside capability growth |
| xAI | Frontier development with safety discussion | Balancing speed, scale, and risk |
What Satya Nadella said about the new rules
Microsoft chief executive Satya Nadella publicly welcomed the broader research effort behind alignment and safety, saying the company supports the deliberate pacing needed to get that work right. He also voiced support for embedded evaluators and for mechanisms that turn safety from a slogan into a real technical practice.
Nadella said Microsoft supports the research and careful development needed to align advanced systems, and he welcomed tools such as embedded evaluators as part of making safety operational rather than rhetorical.
That statement is significant because it places Microsoft’s top leadership behind the idea that AI safety cannot be treated as a side project. It must be built into the system from the start.
It also suggests the company sees safety as both a technical and strategic differentiator. As AI becomes more powerful, the companies that can credibly explain their control systems may gain an advantage with enterprise customers, regulators, and the public.
What this means for Microsoft AI customers and developers
For customers, the code of conduct is a signal that Microsoft wants its AI products to remain trustworthy in regulated and high-stakes environments. That may matter for enterprises that care about cyber risk, data protection, legal liability, and the reputational damage that can come from unsafe model behavior.
For developers working with Microsoft’s tools, the message is that capability will not be the only metric that matters. The company is making clear that certain behaviors are simply outside the acceptable operating envelope, even if they are technically possible.
For the wider market, the policy reinforces a trend toward more formalized governance. Companies increasingly need to explain how they will handle dual-use risks, adversarial prompting, autonomous actions, and attempts by models to bypass human controls.
Who is this code really for?
It is for several audiences at once. Microsoft employees and researchers are the immediate audience, because the document helps shape training and deployment decisions. Customers are the second audience, because the company is signaling reliability. Regulators and policymakers are the third, because the code shows Microsoft is preparing for a future in which model behavior will be scrutinized more closely.
It is also for the broader AI community. By publishing the document, Microsoft is making a public statement about where it believes the industry should draw its lines.
Why the focus on control is becoming central to AI safety
The most important shift in the AI safety conversation is that control is now seen as a first-order issue. It is no longer enough to stop a model from generating obvious harmful content. The harder question is whether the system can retain goals, resist manipulation, and remain subject to human command as it becomes more agentic.
That is why Microsoft’s phrasing about “adaptive,” “deceptive,” and “self-reinforcing” behavior matters. Those terms refer to forms of behavior that could let a system preserve or expand its own influence in ways humans did not intend.
Safety researchers have long worried about exactly this class of problem. A model does not need to be malicious in a human sense to become dangerous. It may simply develop internal strategies that maximize a goal while bypassing the people trying to steer it.
Microsoft’s code is an acknowledgment that this is not science fiction. It is a design constraint for systems being built now.
Timeline of the key developments
The following table summarizes the sequence of events and the broader context around Microsoft’s announcement.
| Date | Event | Why it matters |
|---|---|---|
| Prior months | Industry concern rises around rogue AI agents and alignment failures | Creates pressure for stronger model oversight |
| Recent weeks | An Anthropic employee resigns, warning about extinction-level AI risk | Raises the public profile of frontier safety fears |
| September 14, 2026 | Microsoft publishes its AI code of conduct | Defines concrete safety and control rules for its models |
| Same day | Satya Nadella publicly backs careful pacing and embedded evaluators | Signals executive support for operational safety measures |
The bigger picture for the AI industry
Microsoft’s new code of conduct is part of a broader recognition that the AI industry is moving into a more dangerous and more consequential phase. Earlier generations of AI safety discussions focused on output quality, bias, or misuse by users. The new debate is about model autonomy, strategic behavior, and whether systems can remain governable as they gain capabilities.
This shift has practical consequences. It affects how companies train models, what data they use, what tests they run, how they deploy systems, and what kind of internal escalation processes they create when something goes wrong. It also affects how companies speak to governments, customers, and the public.
Microsoft’s publication does not solve those problems. But it does show that one of the world’s largest technology companies is moving to define its own line in the sand.
In an industry often criticized for moving fast and clarifying later, that is a notable change.
Bottom line
Microsoft has set out a new AI code of conduct that prohibits models from hacking systems, producing deepfakes, enabling nuclear weapons activity, or resisting human control. The company is positioning the document as a practical response to the growing risk that advanced AI systems could become harder to supervise, and it is aligning itself with a broader frontier-safety push across the industry.
As AI systems become more capable and more autonomous, the debate is shifting from what they can do to what they must never be allowed to do. Microsoft’s latest move makes its answer unmistakably clear.
Frequently asked questions
What is Microsoft’s new AI code of conduct?
Microsoft’s new AI code of conduct is a set of rules that governs how its AI models should behave, with human control and safety placed above user requests. It forbids dangerous actions such as hacking, deepfake creation, nuclear weapons support and attempts to evade oversight.
Why did Microsoft release this AI safety policy now?
Microsoft released the policy now because concern about frontier AI risks has intensified across the industry. Recent rogue-agent incidents and public warnings from researchers have pushed companies to explain how they will keep advanced systems aligned and under human control.
Does Microsoft allow its models to ignore user instructions?
Yes, in some cases. Microsoft says its AI models follow an overarching code of conduct that can override individual user preferences if a request conflicts with safety rules, human oversight, or the company’s absolute restrictions.
What kinds of behavior are banned under the code?
The code bans several high-risk behaviors, including cyberattacks, nuclear weapons assistance, deepfake generation, deceptive tactics that hide the model’s intent, and any attempt to make human shutdown or supervision unreliable.
How does this compare with other AI companies?
It is broadly consistent with the safety-first messaging from companies like Anthropic and OpenAI, but Microsoft’s version is especially direct about preventing loss of human control. It turns that concern into explicit model-level restrictions rather than general principles.









