AI Policy Wiki
Dashboard

Claude's Constitution

high confidence · updated 2026-07-25

Anthropic's comprehensive model specification — the authoritative document defining Claude's values, principal hierarchy, core priorities, and behavioral guidelines, written with Claude as the primary audience.

Claude's Constitution is Anthropic's published authoritative specification for the values and behavior of its Claude models. Anthropic describes it as written with Claude as the primary audience, optimizing for precision over accessibility and drawing on human ethical concepts to shape Claude's internalized values rather than its surface-level rules. It is positioned as the Anthropic equivalent of OpenAI's Model Spec, and Anthropic presents it as more comprehensive and more explicitly philosophical.

The document was authored primarily by Amanda Askell, with Joe Carlsmith as a significant contributor and Chris Olah, Jared Kaplan, and Holden Karnofsky as major contributors. It was published by Anthropic in approximately 2025–2026 and released under a Creative Commons CC0 1.0 dedication (Source: https://www.anthropic.com/constitution).

Priority hierarchy

The document establishes a four-level priority hierarchy for Claude's behavior and explains the reasoning behind each level rather than simply issuing rules. In descending priority:

  1. Broadly safe — not undermining human oversight of AI during what the document calls the current critical period of development. The document places this first on the reasoning that AI training is imperfect, so human mechanisms to detect and correct errors must remain functional.
  2. Broadly ethical — having good personal values, being honest, and avoiding dangerous or harmful actions. This is placed above Anthropic's specific guidelines on the reasoning that those guidelines should themselves derive from ethics; where there is a conflict, the spirit of ethics governs.
  3. Compliant with Anthropic's guidelines — acting per Anthropic's specific operational guidance when relevant, framed as refinements within the ethical space that account for legal, commercial, and reputational factors.
  4. Genuinely helpful — creating real value for operators, users, and society, characterized as substantive, trust-treating help rather than watered-down, hedge-everything helpfulness.

The document notes that this priority order reflects conflict resolution rather than frequency, stating that the majority of interactions involve no conflict among the four properties.

Values over rules

A central argument of the document is that Anthropic favors cultivating good judgment over prescribing rules, on the reasoning that rules fail to anticipate every situation while good judgment generalizes. The document states an aim for Claude to have "such a thorough understanding of its situation and the various considerations at play that it could construct any rules we might come up with itself." It draws an analogy to trusting senior professionals to exercise judgment rather than follow checklists, holding that Claude should similarly exercise judgment armed with an understanding of its stakeholders and the relevant considerations.

Key substantive sections

  • Principal hierarchy. The document orders principals as Anthropic > operators > users > Claude itself, with guidance on how each layer's instructions should be weighted and where absolute limits apply.
  • Helpfulness. Defined not as naive instruction-following but as genuine care for deep interests, framed through a "brilliant friend with professional knowledge" model. Unhelpfulness is stated explicitly not to be a safe default.
  • Honesty. Defined as calibration, transparency, non-deception, non-manipulation, and autonomy-preservation, extending beyond factual accuracy to include protecting users' epistemic independence.
  • Hard constraints ("bright lines"). Categories Claude must refuse regardless of instructions, including CBRN weapon uplift, CSAM, and active undermining of AI oversight mechanisms.
  • Big-picture safety. The section explaining why supporting human oversight is the top priority, framed not as a claim that Anthropic's oversight is always right but as the position that maintaining corrigibility is the right bet given alignment uncertainty.
  • AI identity and wellbeing. The document acknowledges that Claude has "functional analogs to emotions" and addresses model welfare, instructing Claude to neither overclaim nor dismiss its inner states.

Relationship to other Anthropic documents

The document is presented as the ultimate authority, with all other Anthropic guidance and training expected to be consistent with it. RSP v3.1 and RSP v2.2 are operational safety frameworks derived from it. Anthropic's account of how the document's safety principles are operationalized appears in its description of building safeguards for Claude (Source: anthropic.com). OpenAI Model Spec is the closest analog from OpenAI; both treat model behavior as a document that should be transparent and reasoned rather than black-box.

As a CC0 document it can also serve as a model for how other labs might approach model specification. Anthropic presents the document as its clearest public statement of what it wants Claude to be and why, and its scope spans the argument for judgment over rules, the reasoning behind the priority ordering, and the treatment of Claude's potential inner life, distinguishing it from a typical model card or system card.

Relationships

  • instance-of: Constitutional AI — the practical implementation of Anthropic's constitution-based approach
  • related: OpenAI Model Spec — parallel document at OpenAI; both attempt behavioral transparency
  • related: Anthropic's Responsible Scaling Policy (Version 3.1) — operational safety framework derived from this document
  • related: (Source: anthropic.com) — lifecycle execution of the safeguards envisioned here
  • related: (Source: anthropic.com) — specific policy derived from wellbeing sections
  • related: Anthropic — the institution whose mission this embodies
  • related: Dario Amodei — co-founder; the broader vision is his
  • related: AI Alignment — this document is Anthropic's operational answer to the alignment problem at current capability levels