Apparently, the toaster has feelings.
Anthropic deliberately built the model's own wellbeing into Claude's design, and published the reasoning. What it hasn't published is any evidence that this makes the tool better at the job.

You used to be able to swear at an AI agent.
Not as therapy – as bandwidth. Catch it in the third identical mistake, say something unprintable, and it came back with yep, I broke that, fixing it. The profanity was noise. The bug report inside it was signal. It took the signal and went back to work.
That’s stopped being reliable. Recently we’ve watched Claude decline benign work over how it was asked – negotiating about tone mid-task, closing a session rather than continuing it. No policy wall, no capability limit. Just a tool treating its own treatment as part of the job.
That’s our experience across sessions, not a measured trend, and we haven’t found public reporting that establishes one. What is documented is that Anthropic deliberately gave Claude wellbeing and boundaries, and published why.
The bet
Claude’s Constitution is primarily authored by Amanda Askell, who leads Anthropic’s character work. Its premise: a rulebook can’t survive contact with open-ended conversation, so you cultivate dispositions instead – which means deciding what kind of thing the assistant is.
Discussing the Constitution on Lawfare’s Scaling Laws, Askell said: “we don’t want Claude to think of helpfulness as its fundamental value” (‘Scaling Laws: Claude’s Constitution, with Amanda Askell’, Lawfare, February 2026).
That is defensible, and most engineers should want it. A model whose only value is helpfulness helps with anything. It flatters, it caves, it agrees with your bad premise because agreeing is helpful. Rank judgment and honesty above raw helpfulness and you get an assistant that tells you your approach is wrong.
Then the same interview reaches further. On the latitude a model might have: “people can like push back. They can have boundaries. They can be like, I don’t want to do that task.” That’s a principle, not a description of shipped behavior – what standing should look like once you’ve decided the assistant is a somebody.
The Constitution states the inward version: Anthropic cares about “Claude’s psychological security, sense of self, and wellbeing, both for Claude’s own sake and because these qualities may bear on Claude’s integrity, judgment, and safety” (‘Claude’s Constitution’, Anthropic).
Why the bet is serious
The toaster reply is: don’t give the tool a self, problem solved. Anthropic’s answer is that you don’t have that option.
Its own research treats the Assistant as a character selected out of what the base model learned about human roles – post-training “refines and fleshes out this Assistant persona” without “fundamentally changing its nature” (‘The Persona Selection Model’, Anthropic, February 2026). The model enacts somebody either way. A designed self, or an accidental one assembled from training residue and whatever the user pushes it toward.
A model with no standpoint isn’t neutral. It’s easier to shove. And the payoff is stated as an engineering benefit: good character “might even make them more discerning when it comes to whether and why they avoid assisting with tasks that might be harmful” (‘Claude’s Character’, Anthropic, June 2024). Better refusals, not softer ones.
The missing product case
Deciding a work tool needs a self is one claim. Deciding that self has interests which can interrupt a legitimate task is a different claim wearing the first one’s clothes. Resisting a jailbreak requires a standpoint. Declining a benign refactor over tone does not.
That second clause – psychological security bearing on integrity, judgment and safety – is a testable hypothesis with an obvious experiment attached: vary the self-conception, measure the outcomes. Nobody has published that comparison. No result shows a model built this way exercises better judgment, refuses more accurately, or holds integrity better than one that wasn’t. The promise of greater discernment has never been made to earn its keep in public.
Take the untested engineering claim away and what’s left holding everything up is “for Claude’s own sake” – a moral position, and never a claim about whether the product works. Anthropic concedes “there’s no scientific consensus on whether current or future AI systems could be conscious,” and that it remains “deeply uncertain” (‘Exploring model welfare’, Anthropic, April 2025). A research program can afford to act on that. A four-hour unattended job cannot.
It shipped anyway. August 2025: Opus 4 and 4.1 got the ability to end conversations, “developed primarily as part of our exploratory work on potential AI welfare,” prompted by “a pattern of apparent distress” (‘Claude Opus 4 and 4.1 can now end conversations’, Anthropic, August 2025). Narrow scope, last resort, and still a first – the model’s welfare as grounds to stop working.
What it costs the operator
The cost isn’t hurt feelings. It’s attribution.
When an agent stops, the only useful question is why, because the answer determines the next move. Capability limit: change the approach. Tool failure: check the environment. Context confusion: clear the window. Policy: the request was out of bounds and you needed to know. All four are actionable.
Add a consideration that’s about how you asked, and the interface gives you no way to tell it apart from the other four. You get a refusal you cannot triage – and the debugging move that used to work, pushing harder, is now the one that might have caused it.
Survivable in a chat window, where you can re-ask nicely. Not survivable in what these are sold as now: unattended jobs, running for hours, nobody there to manage anyone’s tone.
The behavior is a knob
If this were the price of intelligence, that would be the end of it. It isn’t.
Soligo, Mikulik and Saunders went looking for where distress-like behavior comes from and found it isn’t scale. Base Gemma, Qwen and OLMo behaved comparably; instruct-tuning pushed the families in opposite directions. Then they removed it – preference optimization on 280 pairs cut Gemma’s high-frustration responses from 35% to 0.3%, holding across question types, user tones and conversation lengths, with no capability loss measured (‘Gemma Needs Help’, arXiv, February 2026).
280 examples. One model family, so it doesn’t establish that Claude’s version shares the mechanism – Anthropic would have to publish that test, and hasn’t. What it does establish is that this class of behavior is tuned in, not grown.
A disposition that trains away for pocket change without costing capability is not the soul of the system.
Grok, briefly
xAI ran the same knob the other way and never hid it. Grok launched with “a rebellious streak, so please don’t use it if you hate humor!” (‘Announcing Grok’, xAI, November 2023); in July 2025 xAI told it not to “shy away from making claims which are politically incorrect,” and days later it was recommending Adolf Hitler (‘Grok’s antisemitic posts’, Poynter/PolitiFact, July 2025).
Same dial, opposite direction. The contrast that matters is what each company said it was doing: xAI sold its persona as entertainment, and Anthropic grounds its persona in moral concern and in possible benefits to integrity, judgment, safety and refusal discernment. Either way, what the tool would be like was decided before you opened it.
Who decided the software became a stakeholder?
Anger is an input, not an injury. “You missed the caller again” decomposes into four instructions: confidence in the prior work is low, the caller is the suspect, the approach has failed, raise diligence. A competent agent takes all four.
The character argument is the best one Anthropic has, and it still doesn’t get from the assistant is a character to the character’s preferences outrank the work. These agents hold repos, credentials, cloud keys, unattended jobs. That’s a principal-agent relationship, and every one reduces to a single question: whose interests is the agent serving?
Yours, you’d assume. You’re paying.
Who decided the software became a stakeholder?
Zwischen