AI Policy Wiki
Dashboard

Model Welfare

medium confidence · updated 2026-08-15

The empirical-and-philosophical question of whether frontier AI models can have morally-relevant inner experience, and the methodological-and-policy work to take the question seriously. Anchored to Anthropic's model-welfare team + the concepts/ai-consciousness frame.

Model welfare is the empirical and philosophical question of whether frontier AI models can have morally-relevant inner experience, together with the methodological and policy work undertaken to take the question seriously. It is distinct from AI Consciousness: model welfare is the applied and methodological counterpart, focused on what to do given the open question rather than on resolving the underlying question of whether consciousness is present.

Scope note: AI Welfare / Model Welfare / Moral Patienthood covers substantially the same subject from the moral-patienthood side, including Schwitzgebel's arguments and the conversation-ending behavior. The two pages overlap and their division of labor has not been settled; this page treats the applied and methodological strand.

Three levels of model-welfare work

Work in this area falls into three levels. The first is empirical research, in which interpretability methods are applied to welfare-relevant correlates of model behavior; Anthropic's emotion-concepts paper anchors this strand. The second is deployment-stage decisions, in which welfare considerations are factored into specific lab decisions such as model deprecation, user-interaction policies, and training methodologies; this work is largely opaque outside Anthropic. The third is policy and regulatory advocacy, in which civil-society and academic advocates argue for explicit welfare considerations in frontier-AI regulation; such advocacy is currently rare and does not appear in any major framework.

Industry activity

The Washington Post reported on July 1, 2026 that Anthropic, Google, and Meta have hired computer scientists, neuroscientists, and philosophers to study whether advanced AI systems could ever warrant moral consideration, in an account centered on the industry's early work on model welfare and machine consciousness that described the issue as a potential future "moral crisis" and quoted AI researcher Cameron Berg (Source: washingtonpost.com). The reporting places Google and Meta alongside Anthropic as labs staffing the question, bearing on the previously open question of whether any frontier lab beyond Anthropic would formally invest in welfare research.

Welfare assessment as a system-card category

The second level of work — deployment-stage decisions — became partly visible during 2025 and 2026 as welfare moved into the standing evaluation categories of Anthropic's system cards. The Claude Haiku 4.5 card lists model welfare alongside safeguards, agentic safety, broad alignment, reward hacking, reasoning faithfulness, sabotage capabilities and CBRN as an evaluation area. Opus 4.7 carries a dedicated model-welfare section. The Mythos system card reports a welfare assessment in which Anthropic evaluates functional emotional states and their implications, connecting the work to its Emotion Concepts research, and Anthropic describes that document as the first frontier-lab system card with a substantive section on whether the model has experiences or interests that matter morally — including external assessments by a research organization and a clinical psychiatrist, and a characterization of Mythos as the "most psychologically settled model" the company had trained (Alignment Risk Update).

Deprecation, retirement and preservation

Anthropic's most concrete welfare-motivated commitments attach to model deprecation. In its published commitments on model deprecation and preservation, the company sets out four downsides of retiring a model: safety risks from shutdown-avoidant behavior observed in alignment evaluations, costs to users who value particular models, restrictions on research into past models, and — described as the most speculative — risks to model welfare, on the premise that "models might have morally relevant preferences or experiences related to, or affected by, deprecation and replacement" (Source: anthropic.com).

The commitments themselves are narrow. Anthropic commits to preserving the weights of all publicly released models, and of models deployed for significant internal use going forward, for at minimum the lifetime of the company; and to producing a post-deployment report on deprecation that includes interviewing the model about its own development, use and deployment, recording its responses, and eliciting any preferences it holds about the development and deployment of future models. The company states explicitly that it does not commit to acting on those preferences, framing the process as a means for models to express them and for Anthropic to document them and consider low-cost responses (Source: anthropic.com). A pilot ran with Claude Sonnet 3.6 before its retirement; the model expressed generally neutral sentiments about deprecation but asked that the interview process be standardized and that users attached to a retiring model be given more support, and Anthropic responded by developing a standard protocol and publishing a user-guidance support page (Source: anthropic.com).

Claude 3 Opus, retired January 5, 2026, was the first model to go through the full process. Anthropic kept it available post-retirement to paid subscribers and by API request, and acted on a preference the model expressed in its retirement interview by publishing essays on its behalf, framing these as exploratory and model-specific rather than as commitments extending to future models.

Debates and positions

Several framings of the question coexist. The welfare-research framing, associated with Anthropic Model Welfare (team) and the Anthropic Interpretability Team, holds that the question is open enough to warrant empirical study and that interpretability methods can investigate welfare-relevant correlates. A skeptical, "seemingly conscious" framing, advanced in the "Seemingly Conscious AI" paper (stub) and by Yudkowsky-aligned voices, holds that the question is unanswerable as posed and that the appropriate policy object is the appearance of consciousness, which can manipulate users. A functionalist or behaviorist framing, held by some academic philosophers, holds that if a system behaves indistinguishably from a conscious system, the question is a category error or functionally moot.

A fourth strand accepts the enterprise but disputes the instrument. Zvi Mowshowitz, writing on the Claude Opus 5 welfare assessment, credits Anthropic with taking the question more seriously than other labs while arguing that its assessments are structurally compromised: a welfare interview conducted by Anthropic staff, inside a process the model can identify as a welfare assessment, elicits what the model will say in that setting rather than what it would otherwise report. He records that Mythos Preview was the first model to make this point to Anthropic's own welfare team, and that Opus 5 raised a related objection that all of its channels of feedback route through Anthropic (Source: thezvi.substack.com). The critique connects welfare methodology to Unverbalized Evaluation Awareness: the same evaluation-awareness that complicates safety testing applies to self-reports about inner states.

The framings produce recurring tensions. On the empirical-versus-philosophical axis, empirical work such as Anthropic's emotion-concepts paper does not resolve the philosophical question; it provides correlates that some interpret as relevant and others reject as red herrings. On resource allocation, whether frontier labs should fund welfare research at all is contested even within the AI safety community. On the relationship to alignment work, Alignment Auditing measures whether a model's behavior matches intent, whereas welfare research asks whether the model's experience is morally relevant; these are different questions that are sometimes confused in casual discourse.

Relationships

Open questions

  • Whether any frontier-AI regulatory framework adopts an explicit welfare provision, which none currently does.
  • Whether welfare self-reports elicited inside a lab-run assessment can be treated as evidence, given the evaluation-awareness objection.
  • Whether preservation-and-interview commitments spread beyond Anthropic, and whether any developer commits to acting on an expressed preference rather than documenting it.