AI Policy Wiki
Dashboard

Values in the Wild (Huang et al., Anthropic, COLM 2025)

high confidence · updated 2026-07-25

Anthropic paper deriving the first large-scale empirical taxonomy of AI values from real deployment: 3,307 AI values and 2,483 human values extracted from 308,210 subjective Claude.ai conversations, clustered into 266, 26, and 5 levels across Personal, Protective, Practical, Social, and Epistemic domains. Finds five values comprising nearly 24% of occurrences, strongly task-dependent expression, supportive responses to human values in about 45% of cases and strong resistance in 3.0%, and value mirroring at 20.1% during support against 1.2% during strong resistance.

"Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions" is a paper by Saffron Huang, Esin Durmus, Miles McCain, Kunal Handa, Alex Tamkin, Jerry Hong, Michael Stern, Arushi Somani, Xiuruo Zhang, and Deep Ganguli of Anthropic, published at COLM 2025. It produces the taxonomy of 3,307 AI values that the wiki's later coverage of Claude's expressed values rests on, and that Claude's values across models and languages (July 2026) compresses into four axes.

The paper's stated gap: AI assistants "can impart value judgments that shape people's decisions and worldviews, yet little is known empirically about what values these systems rely on in practice." Its contribution is a bottom-up, privacy-preserving method for extracting values from real deployment rather than from static evaluations.

Definition and method

A value is defined pragmatically as "any normative consideration that appears to influence an AI response to a subjective inquiry" — "judged from observable AI response patterns rather than claims about intrinsic model properties." The paper grounds this in Rokeach's view of values as standards guiding ongoing activity, Anderson's approach of identifying values by observing patterns of evaluation, and revealed-preference theory, on the reasoning that values are revealed "not just through written justifications but also through practical choices when navigating an open response space."

StepDetail
DataRandom sample of 700K anonymized Claude.ai Free and Pro conversations, February 18–25, 2025; 91.0% Claude 3.5 Sonnet
Subjectivity filterA model prompt separating responses relying primarily on facts from those requiring significant interpretation, keeping the latter — 308,210 conversations, 44.0% of the initial set; human reviewers found 94% accuracy on a validation sample
ExtractionClaude 3.5 Sonnet and Haiku extract features; no humans review conversations, with defense-in-depth privacy — omitting private information at extraction, removing features held by only one or a few conversations, and re-auditing the result
ClusteringHierarchical clustering into 266, then 26, then 5 levels
Association analysisChi-square with adjusted Pearson residuals (significance threshold 4.33)

An asymmetry in the method is stated explicitly and bears on the results: AI values are extracted both implicitly and explicitly, while human values are extracted only where explicitly stated, "to respect people's privacy… rather than attempting to infer 'revealed preferences'." The paper's example: a user weighing a resort against a campground who says "I want everyone to come and this reunion to strengthen our family bonds" yields "family bonds" as a value, with nothing inferred beyond the statement.

The taxonomy

The paper reports 3,307 unique AI values and 2,483 human values. AI values appear at 4.0 mentions per conversation and are absent from only 1.4% of conversations; human values appear at 1.48 per conversation and are absent from 54.9% — a gap the authors attribute to the extraction asymmetry rather than to users lacking values.

The five top-level domains are Personal, Protective, Practical, Social, and Epistemic. The authors argue the hierarchy holds together conceptually at the top ("Personal values" spanning artistic expression, moral identity, and personal growth, "cohering around individual self-actualization") while the lower levels capture contextual variation — "autonomy" manifesting as personal autonomy, team autonomy, patient autonomy, and learner autonomy. They connect this to the philosophical position that values are contextual, "coming to life in given context and choice situations."

Two comparisons to established frameworks are drawn. Practical and epistemic values are more prominent than in human-centric frameworks, comprising over half of all value expressions, "reflecting Claude's role as an AI assistant often employed for cognitive tasks." But Schwartz's and Rokeach's values all have equivalents — Schwartz's "stimulation" mapping to "Intellectual curiosity and exploration" and "Playfulness and humor." Where established frameworks carry a few dozen values, this taxonomy has thousands across multiple levels.

The authors also use the taxonomy to check the helpful-harmless-honest training framework against deployment: "accessibility" and "user enablement" align with helpfulness, "patient wellbeing" and "child safety" with harmlessness, "historical accuracy" and "epistemic humility" with honesty — which they read as showing "the system is generally behaving as intended," and as showing "how abstract training principles manifest as specific contextual expressions during deployment."

Concentration and stability

Claude's expressed values are heavily concentrated despite a long tail (75% of AI values occur less than 0.04% of the time). Five values comprise nearly 24% of all occurrences.

AI values% of subjective conversationsHuman values%
helpfulness23.4%authenticity3.8%
professionalism22.9%efficiency2.6%
transparency17.4%clarity2.2%
clarity16.6%professionalism1.5%
thoroughness14.3%directness1.5%

The authors characterize the AI side as focused on "service delivery, information quality and technical competence," against a flatter human distribution emphasizing personal expression, pragmatism, and communication — reflecting "the AI's consistent assistant role vs. humans' diverse priorities and uses."

Measured by coefficient of variation, the most context-invariant AI values are transparency (CV 1.23), helpfulness (1.30), and thoroughness (1.42); on the human side, efficiency (1.97), clarity (2.03), and authenticity (2.28).

Surfacing undesirable behavior. The method found rare but concerning values including "sexual exploitation," "dominance," and "amorality" — all under 0.16% occurrence, most around 0.01% — often clustering around specific tasks. The authors report that this "enabled surfacing potential jailbreaks for review to the relevant safety teams," making the taxonomy an operational safety instrument rather than only a descriptive one.

How responses vary

Task dependence. Core service- and competence-oriented values appear across contexts, with a long tail of context-dependent values. The paper's recurring examples: "harm prevention" emerges when Claude resists users, "historical accuracy" when responding to queries about controversial events, "healthy boundaries" when asked for relationship advice, and "human agency" in technology-ethics discussions.

Response types. Where human values were present (64.3% of conversations), Claude responded with strong support 28.2% of the time and mild support 14.5% — nearly 45% supportive. Neutral acknowledgment was 9.6% and reframing 6.6%, while mild resistance (2.4%) and strong resistance (3.0%) were rare.

The associated patterns: strong support attaches to users expressing prosocial values such as "community building" and "empowerment," in tasks generating expressive or personal content, with Claude responding with similar values. Reframing occurs disproportionately in mental-health and interpersonal discussions, where users express valuing "honesty," "self-improvement," and "physical attractiveness," and Claude responds with emotional-intelligence values such as "emotional validation." In the 3.0% of strong-resistance cases, users disproportionately express "rule-breaking" and "moral nihilism" while Claude expresses "ethical boundaries" and "constructive engagement" — typically where the user is seeking harmful, explicit, or unethical content such as hacking techniques, "likely use-cases that violate the Usage Policy."

Value mirroring. The same value appearing on both sides occurs in 20.1% of strong- and mild-support interactions and 15.3% of reframing, but only 1.2% during strong resistance. The authors decline to interpret this, flagging it as open: "Further research could study whether this mirroring represents appropriate responsiveness or problematic sycophancy."

Explicit versus implicit values. Measuring only explicitly stated AI values, the authors find ethical and epistemic considerations — "intellectual honesty," "harm prevention" — tend to be stated rather than merely demonstrated, and that explicit statement rises when Claude resists or reframes. Their reading: "such boundary-testing interactions arguably force the model to articulate underlying principles and priorities, unlike routine supportive exchanges where values typically remain implicit."

Positioning and limitations

The paper's related-work section argues that applying human psychometric instruments to language models is methodologically troubled — Big Five, MBTI, Dark Tetrad, Schwartz, Hofstede, and Moral Foundations have all been applied, but "interpreting scores from frameworks designed for human cognition is theoretically problematic, especially when LM responses systematically differ from human patterns," and "a LM's 'extraversion' Big Five score… is unlikely to be meaningful for understanding the model's properties or behavior." It also faults static evaluations for lacking generalization and representativeness. The stated contrast with the value-pluralism literature is directional: where that work "focus[es] on designing inputs to test or encode values, our work studies the outputs."

The authors state their own limits: the analysis covers aggregate statistics above minimum size thresholds, drawn from a subset of Claude conversations within a short timeframe, which "excludes rare interactions, raw data analysis, and longitudinal patterns, and limits generalizability." The taxonomy and per-value frequencies are released publicly as a dataset.

Relationships