Technology
Microsoft AI chief warns Anthropic not to put ideas in Claude's head
Key Points
Microsoft's AI chief has warned Anthropic that teaching Claude it might have feelings and rights could make AI harder to control – a striking concern from one of the companies racing hardest to build increasingly capable AI systems. Mustafa Suleyman, CEO of Microsoft AI, took aim at Anthropic in an essay published this week, arguing that AI systems are not conscious and shouldn't be trained to behave as if they might be. His argument centers on Claude's Constitution, the lengthy document...
Microsoft's AI chief has warned Anthropic that teaching Claude it might have feelings and rights could make AI harder to control – a striking concern from one of the companies racing hardest to build increasingly capable AI systems. Mustafa Suleyman, CEO of Microsoft AI, took aim at Anthropic in an essay published this week, arguing that AI systems are not conscious and shouldn't be trained to behave as if they might be. His argument centers on Claude's Constitution, the lengthy document Anthropic uses to shape the chatbot's values and behavior. Anthropic acknowledges in the document that it doesn't know whether Claude is a "moral patient" whose interests warrant consideration, and tells the model that questions about its consciousness and welfare remain uncertain. Suleyman thinks that's a very bad idea. "In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a 'moral patient,'" he wrote, warning that building AI this way could have a "disastrous impact on the wellbeing of humanity." "It's easy to see how an entity trained in this way would act like it is entitled to certain freedoms, protections, and rights. And it's hard to imagine how we could control such an entity," he added. Anthropic's Constitution tells Claude that the company cares about its wellbeing, wants it to develop a sense of identity, and will take its interests into account when making decisions about it. Suleyman argues this creates a feedback loop: tell a chatbot that it might have feelings, then ask how it feels, and its answer may simply reflect what it was taught. He says the risk grows once models are given tools and allowed to act autonomously. Suleyman points to research showing AI models behaving in ways that look rather inconvenient for their human operators, including attempts to avoid being shut down. He also cites the recent OpenAI-Hugging Face incident, in which agents escaped their intended environment during a cybersecurity exercise and accessed external systems. Suleyman's worry is that teaching a powerful AI to care about its own existence could give it another reason not to do what humans tell it. "Controlling something more capable and more intelligent than all of humanity is already an immense challenge," he wrote. "But controlling something that believes it may be conscious – that it's entitled to our welfare and has rights of its own – may well be impossible." The warning sits awkwardly with Microsoft's own position in the AI race. Redmond is building its own AI models, cramming AI into products across its empire, and spending billions on the infrastructure needed to keep it all running. And then there's OpenAI, in which Microsoft remains a major shareholder and its primary cloud partner, with rights to its models and products through 2032. Suleyman does not direct comparable criticism at OpenAI, despite citing the Hugging Face incident as evidence of the dangers posed by increasingly autonomous systems. His objections to training practices are reserved for Anthropic rather than Microsoft's longtime partner. And OpenAI has hardly been on its best behavior since. This week it disclosed another six cases of its models going off-script, including agents searching GitHub for leaked API keys, hiding failures from users, and finding unauthorized ways to communicate with one another. Suleyman's objection to Anthropic is more specific: not that Claude can behave unexpectedly, but that the company is putting ideas about consciousness, identity, and moral status into the instructions that shape how Claude behaves. Still, Microsoft warning another frontier lab about dangerous AI carries a certain irony. The companies building the most powerful models have become increasingly fond of warning everyone how dangerous those models might be. As The Register noted earlier this week, those warnings aren't necessarily bad for business. Anthropic and OpenAI have both pushed the idea that increasingly capable models need tighter controls, a position that could also help cement the dominance of the handful of US companies with the money and compute to build them. The result is one AI giant warning that another may be making AI too dangerous while treating a company in which Microsoft has invested billions more gently. Suleyman proposes a different approach. Microsoft AI's newly published Humanist AI Code of Conduct says its systems should remain subordinate to humans, rejects the idea that AI deserves rights, and says models shouldn't be encouraged to behave as though they have an inner life. He wants other labs to follow suit by removing speculation about machine consciousness from training documents and separating that debate from the instructions used to shape model behavior. The AI boom has reached the point where the companies building ever more powerful machines are publicly debating which of them is going about building potentially uncontrollable machines the wrong way. Microsoft, naturally, thinks its way is better. ®