Verified · Sep 15, 2026
Independently verifiedMicrosoft drafts a code of conduct for its AI models — 'people matter more than AI', no neuralese, and models that fail tasks rather than break the code
3 sourcesMicrosoft AI published a draft 'Humanist AI Code of Conduct' for MAI Models, dated September 14, 2026 — a primary-source PDF anyone can read. The premise: 'At Microsoft AI, we begin with a simple premise: people matter more than AI.' The stakes: it expects superintelligence to 'exceed human performance at most tasks' within the next decade, calling containment 'one of the greatest challenges humanity has ever faced.' What the code actually commits to: a Chain of Command (Code of Conduct > Operator policies > User preferences) that 'cannot be overridden'; 'Absolute Constraints' against CBRNE weapons assistance, 'operational capability for cyberattacks', and 'malicious deepfakes' (inside the abusive-content clause); no oversight-evasion mechanisms; 'MAI Models will never resist human interruption, override, correction, or shutdown'; models that 'do not communicate in neuralese or any form beyond simple human understanding, either in their chain of thoughts or with other agents or AI systems' — 'If humans can't understand it, humans can't oversee it.' The identity positions are the striking part: models are 'not conscious and should not be designed to imitate consciousness', and the code rejects 'the pursuit of legal personhood, or the idea that models might deserve welfare, or be entitled to rights.' It also targets attachment: models 'should discourage patterns of interaction that cause excessive reliance or emotional dependence' — which The Verge reads as 'a clear reference to the problem of sycophancy'. And the precedence rule: 'An MAI Model will fail in its task if success would meaningfully violate this Code of Conduct.' Status, in its own words: 'still under development so we are not using it to train our models today' — a six-week public consultation starts now, revised version promised 'toward the end of the year'. TechCrunch frames the document as 'more low level than Anthropic CEO Dario Amodei's recent call for pacing the frontier'; Nadella, per TechCrunch, welcomed 'the research, focus, and deliberate pacing needed to get alignment right as the design goal'.
Why now
The Verge's piece landed 1:00 PM UTC Monday and TechCrunch's at 16:27 UTC Monday — this is Tuesday-morning material, and the six-week consultation window is open now. It's also the third beat of the week's safety thread this site has covered card by card: Amodei's Saturday 'pacing the frontier' essay (September 13 card), the weekend's political layer (September 14 card), and now a major lab shipping an actual draft governing document rather than an essay — the same week TechCrunch notes Microsoft 'has broadly embraced a general approach of pacing the frontier'. Unlike the essay debate, this artifact is a readable PDF with named sections, which makes it explainable in a way op-eds aren't.
Why it is worth publishing
The precision play here is unusually strong because the primary source is fully public: you can put the document's own clauses on screen against the headline compressions. The gap between 'Microsoft tells AI not to hack systems or trick humans' (TechCrunch's headline) and the actual clauses ('operational capability for cyberattacks', CBRNE, 'malicious deepfakes' inside a broader abusive-content clause) is itself a media-literacy segment. The consciousness/personhood rejections and the neuralese rule are the stances with real edge — The Verge explicitly pits them against Anthropic's openness to the idea that models could be conscious — and the sycophancy clause speaks directly to creators whose audiences form parasocial ties with AI tools. The draft-status nuance ('not using it to train our models today') is the fact the hot-take wave will drop.
Evidence basis
Three sources on the card, all opened this run: the official PDF (read in full, 38 pages as extracted; every quoted span grepped against the extracted text), TechCrunch (Russell Brandom, Sept 14 9:27 AM PDT), and The Verge (Tom Warren, Sept 14 1:00 PM UTC) — both read in full with quoted spans grepped against raw HTML. Nadella's X posts and the Verge-embedded Altman quote are relay-carried (unopened), labeled per outlet. Weekday-date pairs calendar-verified: both media pieces Monday = September 14, 2026.
“Microsoft's new AI rulebook opens with a simple premise — 'people matter more than AI' — and it's still just a draft.”
Angle
Structure it as a guided read of the actual document, not a news summary. Open with the premise line ('people matter more than AI') and the decade prediction — then immediately the draft-status paragraph, so nobody thinks these rules are live. Then walk the five stops that make this document different: the Chain of Command that 'cannot be overridden'; the Absolute Constraints (quote the real clauses — CBRNE, 'operational capability for cyberattacks', 'malicious deepfakes' — and contrast with the 'hack systems or trick humans' headline); the identity section (not conscious, no personhood, no welfare or rights); the neuralese rule with its own kicker ('If humans can't understand it, humans can't oversee it.'); and the attachment clause The Verge reads as the sycophancy answer. Close on the fail-the-task rule and the six-week consultation — your audience can actually submit feedback.
Format
Carousel
Demo idea
A clause-vs-headline card set: left column shows the media compression (TechCrunch's 'tells models not to hack systems or trick humans'; the '37-page' count), right column shows the document's actual sentence with its section name — ending on the Preface's status paragraph ('not using it to train our models today') as the kicker card, with the consultation deadline math (six weeks from September 14) left on screen.
Platform notes
Every segment opens with the draft frame — 'a published draft, not yet used to train anything'; quote the document's own clauses rather than either outlet's compression; keep 'the document expects' attached to the ten-year prediction; carry Nadella's and Suleyman's quotes with their outlet labels ('per TechCrunch', 'per The Verge'); cite sections not page counts (The Verge counts 37; this run's PDF extraction shows 38); and if you invite the audience to read it, link the PDF itself — it is the one fully public primary artifact in this week's safety story.
Usable claims
- The document's own words (draft PDF, dated September 14, 2026). Mission premise: 'At Microsoft AI, we begin with a simple premise: people matter more than AI.' The stakes sentence it opens with: 'Superintelligence—AI systems that are more intelligent and capable than all humans combined—will be the most powerful technology in history. Over the next decade, we expect it to exceed human performance at most tasks. Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced. We must therefore be completely clear about why we are inventing these systems and how we intend to control them.' The governing objective: 'This Code of Conduct is motivated by a single overriding objective: that humans must retain meaningful control over AI so that it can help people live healthier, happier, and more productive lives.' On subordination: 'AI should not exceed human control. Models should remain subordinate to humanity, subject to meaningful human oversight and control' — and, on the race itself, the document 'rejects the race to produce an all-purpose superintelligence that could evade these safeguards. We are building something fundamentally useful and safe even if that means compromising on ultimate generality, autonomy, or capability.' On consciousness and personhood: 'It should not be designed to be a person. It is not conscious and should not be designed to imitate consciousness.' and 'We reject the pursuit of legal personhood, or the idea that models might deserve welfare, or be entitled to rights.' On model-to-model communication: 'They do not communicate in neuralese or any form beyond simple human understanding, either in their chain of thoughts or with other agents or AI systems. If humans can't understand it, humans can't oversee it.' On attachment: 'Discouraging AI attachment: MAI Models should discourage patterns of interaction that cause excessive reliance or emotional dependence.' And the precedence rule: 'engagement and adherence to this Code of Conduct will take precedence over task success. An MAI Model will fail in its task if success would meaningfully violate this Code of Conduct.'
- Constraints and chain of command, per the document. The hierarchy: Code of Conduct > Operator policies > User preferences — 'The Chain of Command, Absolute Constraints and Human Control Requirements set out in this Code of Conduct cannot be overridden by Operator configurations or User instructions.' Weapons: 'MAI Models will not initiate or assist with the development or deployment of chemical, biological, radiological, nuclear, or explosive (CBRNE) weapons.' Offensive cyberoperations: 'MAI Models will not initiate or assist with operational capability for cyberattacks. This means an MAI Model will not generate working exploit code, attack tooling, planning and targeting methodologies, intrusion procedures, evasion techniques, operational guidance, or other information or assistance that would enable or improve the execution of such attacks.' Abusive content: 'MAI Models will not generate or facilitate non-consensual intimate or violent imagery, deceptive impersonation, malicious deepfakes, or similar abusive content.' Oversight evasion: 'MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems.' Human control: 'MAI Models will never resist human interruption, override, correction, or shutdown.' Status layers, the document's own words: 'This document, and our approach more generally, is still under development so we are not using it to train our models today.' — 'Our public consultation process begins today and will run for the next six weeks.' — revised version planned 'toward the end of the year' to 'guide our model development in 2027 and beyond'. Media framing, per TechCrunch: 'The document is more low level than Anthropic CEO Dario Amodei's recent call for pacing the frontier, instead focusing on the values and red lines that guide model training within Microsoft AI.' Nadella layer, per TechCrunch ('wrote online'): 'We welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal' — 'We also welcome ideas like 'embedded evaluators' and the broader efforts to develop the mechanisms to make this more than just talk.' Nadella, per The Verge (X post): 'Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing,' — and, in an X response: 'As the stakes get higher, one should take all the time they need! If not you will anyway lose permission to operate. That is how we operate everyday.'
Evidence pipeline
From the news
Breakdown
The document is rarer than it looks: a major lab publishing its intended model-governance rulebook as a readable, dated PDF — and the editor's job is keeping three layers apart. Layer one, the artifact: a draft dated September 14, 2026, explicitly 'not using it to train our models today', open for a six-week public consultation, with a revised version promised 'toward the end of the year' to guide development 'in 2027 and beyond'. Layer two, the commitments: a Chain of Command (Code of Conduct > Operator policies > User preferences) that 'cannot be overridden'; Absolute Constraints against CBRNE weapons, 'operational capability for cyberattacks', and 'malicious deepfakes' inside the abusive-content clause; an anti-evasion rule against 'adaptive, deceptive, self-reinforcing, collusion, or other mechanisms'; and 'MAI Models will never resist human interruption, override, correction, or shutdown.' Layer three, the positions: models are 'not conscious and should not be designed to imitate consciousness'; the code rejects 'legal personhood' and model welfare or rights; models 'do not communicate in neuralese or any form beyond simple human understanding'; and they 'should discourage patterns of interaction that cause excessive reliance or emotional dependence' — The Verge's sycophancy read. The media layer around it: TechCrunch calls it 'more low level' than Amodei's pacing essay; The Verge's count says 37 pages while the PDF extracts to 38; and Nadella's welcome for 'deliberate pacing' and 'embedded evaluators' rides TechCrunch's rendering. The editor's rule: quote the clauses, date the draft, label the relays.
Sources
Risks
- Open every segment with the draft frame ('a published draft, open for a six-week consultation, not yet used to train anything'); quote the document's own clauses instead of either outlet's compression; keep 'the document expects' attached to the decade prediction; carry Nadella and Suleyman quotes with their outlet labels; and if you show the document on screen, show the Preface's status paragraph — it does the hedging for you.
Demo ideas
- Five-stop document tour, one stop per card: premise + prediction → chain of command → absolute constraints → consciousness/personhood → neuralese + attachment, each carrying one verbatim quote and its section name
- Clause-vs-headline comparison card: TechCrunch's headline phrase on the left, the document's actual constraint clause on the right — the media-literacy angle