The Anthropic Model Welfare team is an internal research program at Anthropic focused on model-welfare questions: whether AI systems can have morally-relevant inner experience, what empirical methods would test the question, and what deployment-stage welfare considerations follow if the answer is non-trivial. Its work treats welfare as an open empirical question and serves as a reference point for the welfare-research framing of AI Consciousness, in contrast to the more skeptical "seemingly conscious AI" framing.
Activities
The team's principal published output is the emotion-concepts paper (Emotion Concepts and their Function in a Large Language Model), which uses interpretability methods to study Claude's internal-state correlates of emotion-like patterns. According to public Anthropic communications, the team also has input on welfare-relevant deployment decisions, such as model-deprecation choices that consider potential welfare implications. Anthropic's broader interpretability program (Anthropic Interpretability Team) overlaps with this work; the welfare team's distinct contribution is the question-framing rather than the methods.
Positions and debates
The team frames model welfare as an open empirical question. Critics, including many other frontier labs, reject the framing as philosophically confused. Whether frontier labs should fund welfare research at all is contested even within the safety community.
The team's focus is distinguished from alignment auditing, which measures whether a model's behavior matches intent; welfare research instead asks whether the model's experience is morally relevant. The two are different questions that are sometimes conflated.
Relationships
- related: AI Consciousness, Model Welfare, Alignment Auditing.
- related: Emotion Concepts and their Function in a Large Language Model, Anthropic Interpretability Team.
- related: Anthropic (parent).
Sources
Stub created 2026-05-11.