AI Policy Wiki
Dashboard

Anthropic Model Welfare (team)

medium confidence · updated 2026-06-06

Anthropic's internal team / research program focused on model-welfare questions — whether AI systems can have morally-relevant inner experience, what experiments would test the question, and what deployment-stage welfare considerations follow. Anchored to the emotion-concepts paper and the broader concepts/ai-consciousness question.

The Anthropic Model Welfare team is an internal research program at Anthropic focused on model-welfare questions: whether AI systems can have morally-relevant inner experience, what empirical methods would test the question, and what deployment-stage welfare considerations follow if the answer is non-trivial. Its work treats welfare as an open empirical question and serves as a reference point for the welfare-research framing of AI Consciousness, in contrast to the more skeptical "seemingly conscious AI" framing.

Activities

The team's principal published output is the emotion-concepts paper (Emotion Concepts and their Function in a Large Language Model), which uses interpretability methods to study Claude's internal-state correlates of emotion-like patterns. According to public Anthropic communications, the team also has input on welfare-relevant deployment decisions, such as model-deprecation choices that consider potential welfare implications. Anthropic's broader interpretability program (Anthropic Interpretability Team) overlaps with this work; the welfare team's distinct contribution is the question-framing rather than the methods.

Positions and debates

The team frames model welfare as an open empirical question. Critics, including many other frontier labs, reject the framing as philosophically confused. Whether frontier labs should fund welfare research at all is contested even within the safety community.

The team's focus is distinguished from alignment auditing, which measures whether a model's behavior matches intent; welfare research instead asks whether the model's experience is morally relevant. The two are different questions that are sometimes conflated.

Relationships

Sources

Stub created 2026-05-11.