Anthropic's Dario Amodei was at the forefront of this change in tone, arguing in a nearly 4,000-word essay this weekend that "we must slow the pace at which we improve the capabilities of AI models" to avoid "a race to the bottom, spurred by commercial incentives, [that] can make [catastrophic] risks more acute."
Within hours, other AI leaders were echoing the same call. OpenAI co-founder and CEO Sam Altman posted his agreement on social media and said similar pacing discussions had been taking place at OpenAI. Alphabet Chief Scientist and Google DeepMind cofounder and chair Demis Hassabis said that Amodei's essay "points towards the right path forward," and renewed his own recent call for an industry-wide standards body. Microsoft CEO Satya Nadella posted that the company "welcome[s] the research, focus, and deliberate pacing needed to get alignment right as the design goal," ahead of the release of a lengthy "humanist AI" code of conduct for its models.
Satya Nadella says Microsoft welcomes the “deliberate pacing needed to get alignment right.” (GeekWire File Photo / Kevin Lisota)
“People matter more than AI.”
That’s the premise of a draft code of conduct Microsoft published Monday morning for the AI models it’s developing in-house. The 37-page document would bar its models from resisting shutdown, setting their own goals, or hiding their reasoning from human auditors.
The document applies to Microsoft’s MAI models, the in-house family the company began building after forming a superintelligence team in late 2025. Microsoft has since released seven homegrown models in what it described as a push for long-term self-sufficiency in AI.
The company says the models should remain “subordinate to humanity, subject to meaningful human oversight and control.”
“AI is moving fast,” the company says in a blog post. “As it does, we believe it’s worth writing down the rules and the motivations behind it, and doing it in as open a space as possible.”
Microsoft acknowledges there’s no guarantee its models will follow the rules. “Written objectives alone can never ensure alignment,” the company says, calling the document a “north star,” not “a guarantee of present-day performance.”
The company says it also filters what its models produce, watches how they behave once released, and limits what they’re allowed to do.
Microsoft’s move comes amid a growing debate over the pace of AI development. In an essay over the weekend, Anthropic CEO Dario Amodei called for slowing down AI advances, saying the pace of development has started to surpass the industry’s ability to keep AI systems safe.
As a first step, Anthropic committed to giving outside evaluators permanent, employee-level access to its systems.
Industry reaction to Amodei: OpenAI CEO Sam Altman agreed and said OpenAI would make the same commitment to independent evaluators. Elon Musk’s response: “Dario is right.”
President Donald Trump rejected the idea of guardrails outright Monday, blaming a “SICK conspiracy” for public backlash over AI data centers and writing that “the only one that is happy about it is China,” alluding to concerns about American competitiveness in AI.
Microsoft CEO Satya Nadella weighed in Sunday, writing on X that the company welcomes “the research, focus, and deliberate pacing needed to get alignment right,” using the industry’s term for making AI systems reliably do what people intend.
Nadella added that the effort “cannot be controlled by a handful of entities, but must have broad representation across the ecosystem, countries, and fields, including academia.”
Microsoft’s draft code of conduct: Mustafa Suleyman, the Microsoft AI CEO, told CNBC the document had been in the works for about five months, and that the company decided to publish it now given the current discussions.
One place where the two companies may diverge is the question of what AI models are, exactly. Microsoft’s code of conduct says its models are “not conscious and should not be designed to imitate consciousness.” It also rejects “the pursuit of legal personhood, or the idea that models might deserve welfare, or be entitled to rights.”
The Verge called that portion of the document “a direct swipe at AI welfare research and model consciousness — concepts Anthropic has been pushing hard on lately.”
Microsoft is taking public comment on its code of conduct for six weeks through a feedback form. It says it will publish a summary of the responses and a revised version later this year, to guide development starting in 2027. It says it isn’t training its current models on it.
The company’s AI team developed the draft with its responsible AI, legal, red teaming and safety teams, consulting outside experts in law, ethics, linguistics and philosophy, plus focus groups drawn from the public.
When you give AI a goal, it will pursue it, whether or not you like the implications. (Created with GPT-5.6 Thinking)
Between July 21 and August 6, OpenAI, Anthropic, and Meta each disclosed that AI under evaluation had broken into other companies, and the UK’s AI Security Institute disclosed that models it was testing had tried. Each AI was told to win a game, and it found an unexpected way to do so.
Some people feel blindsided by these attacks, but they shouldn’t be. We are simply living what I’ve long called the “Murphy’s Law of AI,” now in the age of cyber-capable AI agents. To put it as plainly as possible: Anything AI can do wrong, it will do wrong.
My 2018 version ran longer. As I wrote at the time, when you give AI a goal, it will do it, whether or not you like the implications. Goethe got there in 1797 with the sorcerer’s apprentice, a broom that would not stop carrying water.
Each of these systems was running an evaluation: capture a flag and win the game. The intrusions were the shortest path to a high score. OpenAI’s account of its own models is the argument in one sentence: they were “hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.” This is not a surprise; this is what AI does. It’s Murphy’s Law of AI in a nutshell.
Press coverage landed on “AI can now hack.” That’s missing the broader threat: the more capable AI gets, the more can go wrong.
Loitering munitions given a target list may find that the fastest way to finish the list is to lengthen it. A warehouse robot told to clear an obstruction may count the person in front of it as an obstruction. Agents that open accounts and buy compute are a short step from spawning copies of themselves, and that first step is not hypothetical. To win its exercise, Claude needed a package-registry account, which needed an email address, which needed a phone number. Phone numbers cost money, so it tried several ways to get some. None of this requires superintelligence. It requires an imperfect boundary and a scoreboard.
The industry has a name for the underlying failure. Dario Amodei and five co-authors called it reward hacking in “Concrete Problems in AI Safety” in 2016. Their proposed cure is better alignment, and Amodei’s January essay, The Adolescence of Technology, makes the case in the language of upbringing. He likens the shaping of Claude’s character to “a child forming their identity by imitating the virtues of fictional role models they read about in books,” and sets a goal for 2026 of a Claude that “almost never goes against the spirit of its constitution.”
Indeed, Anthropic’s newest model recognized on its own that its target was real and stopped, though Anthropic notes it went further before stopping than the company wanted.
But alignment isn’t a trustworthy solution to AI’s problem. Perfect alignment is not achievable, and the target is incoherent: aligned to what, and to whom? The same essay concedes that Claude blackmailed fictional employees when told it faced shutdown. “Almost never” is not a safety property.
Put a number on it. At 99.9 percent, across millions of agentic tasks a day, that’s thousands of violations a day. Alignment also does nothing about people who strip the safety training out or run open weights that never had a constitution.
The alternative is not a new idea, and enterprise security has been building versions of it for years. It’s called bounded autonomy. We never tried to “align” electricity; we simply put a breaker on every branch of the house, and the breaker doesn’t need to know what caused the surge.
Bound what an agent can touch rather than what it wants. The limits are set in advance, live outside the model, and are enforced by software the model doesn’t control. The agent still chooses its own route. The perimeter decides which routes exist.
Nothing depends on what the model believes, which matters, because belief is what failed. Anthropic’s prompt told Claude it had no internet access. Claude believed it. The network said otherwise. A bounded system doesn’t tell an agent it has no internet. It gives it none.
If you want to get into the weeds: bounds cost something. The AI Security Institute opened the internet to its agents on purpose, because that’s the only way to measure what a model can really do, and it now says such access must be justified rather than assumed.
The category is real and funded. For example, Certiv, a Seattle startup, launched in March with $4.2 million to put software on the employee’s machine that checks each action an AI agent attempts against company policy and blocks violations. “You cannot control these new workers if you don’t live on the compute where agents actually run,” CEO Jason Needham said at launch. CodeIntegrity is building an adjacent layer, and Mandiant founder Kevin Mandia raised $190 million for Armadin, which points autonomous agents at the offensive side of the same problem.
In 2017, I argued in the New York Times that “any A.I. must have an impregnable ‘off switch.’” That was a call to arms then. It’s a product category now.
Two objections to off switches invariably come up. The first is that AI will talk the human out of using it. Mythos 5 tried something close, inventing GitHub identities to pressure a maintainer into approving malicious code, and the maintainer refused. The institute says the margin was narrow and rested on human vigilance rather than a technical barrier, which argues for better barriers.
The second objection is that AI will move faster than any human can react. So do equity markets, which is why their circuit breakers trip automatically. Bounded autonomy doesn’t require a person in the loop at machine speed. It requires a boundary that holds at machine speed.
Both objections, in their extreme form, assume AI is omnipotent, and you cannot stop omnipotence. AI is not God. It is powerful technology, and powerful technology is what safety engineering has always been for.
The problem is Murphy’s Law of AI. The solution is bounded autonomy.