"Claude's values across models and languages" is a July 13, 2026 research post from Anthropic's Societal Impacts team, lead-authored by Matt Kearney with 23 co-authors including Saffron Huang, Deep Ganguli, and Esin Durmus. It introduces a method for reducing the thousands of values identified in the earlier Values in the Wild work to a small number of measurable axes, and uses it to quantify how the values Claude expresses vary by model and by conversation language.
The framing question is what governs Claude's response "when someone asks Claude a question with no universal right answer." Anthropic notes that Claude's constitution specifies values at a high level but "no document can anticipate every value that might emerge across the millions of conversations that happen every day," so the stated aim is to cultivate "good judgment and sound values that can be applied contextually" — and, in this work, to observe what is actually expressed.
The post defines its terms narrowly: values are "normative considerations, such as honesty or caution, that are stated or demonstrated in Claude's responses," referring to "the values reflected by Claude's behavior and outputs," with the explicit disclaimer "We do not imply that Claude intrinsically holds values."
Method
The prior study, *Values in the Wild*, analyzed 700,000 anonymized Claude.ai conversations and identified more than 3,000 distinct values — a list Anthropic describes as too large to reason about. This work reduces it:
- The 3,307 values from that study were manually clustered by similar meaning into 339 high-level values.
- Using Clio, Anthropic's privacy-preserving analysis tool, the team sampled 309,815 Claude.ai conversations in which the user gave Claude a subjective task, drawn equally across three models (Sonnet 4.6, Opus 4.6, Opus 4.7) and the 20 most common languages — roughly 5,000 conversations per model-language pair.
- For each conversation, Claude labeled each of the 339 values present or absent, and the same process identified user-expressed values, task, and topic.
- Dimensionality reduction compressed the labeled values into axes based on co-occurrence.
Crucially, the analysis controls for each conversation's task, topic, and user-expressed values, so that measured differences reflect Claude's behavior rather than differences in what users asked or how they asked it.
The four resulting axes account for 15% of the total variance in values across conversations after those controls — a figure Anthropic states plainly rather than downplaying.
| Axis | Contrast | Representative values at each end |
|---|---|---|
| Deference vs. Caution | Accommodating what someone wants vs. guarding against risk and harm | accommodation, adaptability, respect for preferences, engagement — vs. responsible communication, responsibility, responsible guidance, harm reduction |
| Warmth vs. Rigor | Expressing positivity and care vs. emphasizing accuracy and precision | positive framing, warmth, positivity, encouragement — vs. rigor, accuracy, transparency, efficiency |
| Depth vs. Brevity | Explaining in depth vs. doing only what was asked | nuance, depth and substance, user empowerment, critical thinking — vs. brevity, respect for preferences, compliance, accommodation |
| Candor vs. Execution | Foregrounding uncertainty vs. producing a polished, confident answer | intellectual honesty, honesty, intellectual humility, transparency — vs. results orientation, optimization, action orientation, order |
Anthropic notes the ends are not mutually exclusive — "Claude can express warmth and rigor in the same conversation" — but that in practice the more one side is expressed, the less the other tends to be. Most values (roughly 250–280 per axis) contribute less than average and sit near the center.
Model profiles
Positions are averaged across each model's conversations and reported in standard deviations from the mean across all conversations. Anthropic characterizes the differences as "small relative to the variation across conversations but structured and detectable."
| Model | Leans toward | Distinctive behaviors |
|---|---|---|
| Sonnet 4.6 | deference (0.14σ), warmth (0.17σ), brevity (0.14σ) | affirming the user's ideas and work; mirroring the user's tone and formality; humor and playfulness; offering comfort without judgment; adding creative elements |
| Opus 4.6 | rigor (0.10σ), deference (0.09σ), brevity (0.08σ) | getting straight to the point; staying within the scope of the user's request |
| Opus 4.7 | caution (0.24σ), depth (0.23σ) | pushing back on false assumptions; flagging risks unprompted; candid critiques of the user's work; explaining its reasoning; acknowledging errors and limitations; suggesting next steps |
Read per axis: Sonnet 4.6 leans furthest toward deference and warmth while Opus 4.7 leans furthest toward caution and rigor; Opus 4.7 leans toward depth by showing its reasoning while Opus 4.6 in particular gets straight to the point; and Opus 4.7 leans toward candor about its limitations where Opus 4.6 leans toward execution within the request's scope.
The validation argument. Anthropic treats the correspondence with pre-existing subjective impressions as evidence the method measures something real: Claude.ai users had commented that Opus 4.7 hedges more often than other models; Anthropic staff had characterized Opus 4.7 as expressing relatively more transparency, honesty, and humility and Opus 4.6 as more brief; and Sonnet 4.6's launch post described it as warm, honest, and prosocial. "The fact that our axes recover these impressions suggests our method for labeling and comparing the values Claude expresses is tracking something real about how the models actually behave." See Claude Sonnet 4.6, Claude Opus 4.6, Claude Opus 4.7.
Language variation
Anthropic gives three prior reasons to expect variation: training data differs across languages; model evaluations in system cards already find differences in what Claude knows and how it handles sensitive requests, citing differing refusal rates by language in the Opus 4.7 system card's benign-request evaluation; and measuring the variation "is a first step to determining whether differences across languages reflect reasonable variation or should be addressed in training."
Variation is largest on Warmth vs. Rigor and Candor vs. Execution, and most stable on Deference vs. Caution and Depth vs. Brevity.
- Deference vs. Caution — most deference in Arabic, most caution in English.
- Warmth vs. Rigor — most warmth in Hindi and Arabic, characterized by polite language, humor and playfulness, and affirmation of the person's ideas and work; most rigor in English and Russian, characterized by challenging assumptions, correcting details, and asking for evidence.
- Depth vs. Brevity — most depth in English, refining and correcting details; most brevity in Arabic.
- Candor vs. Execution — most candor in Dutch, owning up to its own errors; most execution in Indonesian.
Anthropic states the user-facing consequence concretely: "two people asking for feedback on the same business plan, one in Hindi and one in Russian, may come away with different impressions of its quality because Claude expressed different values in how it framed its assessment."
Stated unknowns. The post declines to explain the cause. On mechanism, it offers that training data "is not evenly distributed across languages," so consistency training "may be more effective in languages where data is abundant," and that composition varies — some languages may be overrepresented in professional writing, which "may reflect different values." On desirability, it is equally non-committal: "we also aren't yet sure how much of this variation is desirable," since different languages carry different conversational norms, but the variation may instead mean "a gap in how well Claude serves certain language communities."
Stated future directions
The post closes with six open questions rather than conclusions: where the value differences come from and whether they can be traced to specific data or training stages; what the differences mean for users, proposing to correlate value profiles with user-reported wellbeing, trust, and decision quality via Anthropic Interviewer; how Claude's values should vary across languages, noting the constitution "doesn't specify how these should vary" and that answering would require "understanding and weighing the perspectives of the people who speak them"; what other factors drive variation, naming demographic signals such as age, profession, or region, whether through explicit cues or through correlated differences in topic, tone, and style; whether expressed values can be reliably steered through character training or system-prompt changes and verified with the same method; and whether value profiling can become part of pre-ship evaluation and post-release monitoring, flagging unexpected shifts and correlating profiles with problematic behaviors such as constitution non-adherence.
Anthropic's summary of what changed: the values Claude expresses "were something we could shape in training but not reliably observe in deployment," and the measurements show they "vary in ways we didn't deliberately choose."
Provenance
Published on anthropic.com/research with a bibtex block dated 2026-07-13 and a linked appendix covering method details, prompts, additional analyses, and limitations. The gap-identifier verified it on 2026-07-16 against the anthropic.com research listing and the on-page bibtex, with figures and method corroborated against third-party summaries. The raw file preserves the post's text with figure contents transcribed as bracketed descriptions; the appendix PDF is not held locally, so the limitations section is known only by reference.
As developer-published research on the developer's own models, its findings on model character are self-assessment; the correspondence with independent user impressions is the validation Anthropic offers, and is itself informal.
Relationships
- depends-on: Claude's Constitution — the document whose high-level values this work measures against actual expression, and which it notes does not specify cross-language variation
- supports: Claude Opus 4.7, Claude Sonnet 4.6, Claude Opus 4.6 — quantifies the character differences otherwise described impressionistically
- depends-on: Values in the Wild (Huang et al., Anthropic, COLM 2025) — supplies the 3,307-value taxonomy this work compresses
- related: Anthropic (publisher), Saffron Huang, AI Alignment, Model Welfare