AI Policy Wiki
Dashboard

Saffron Huang

medium confidence · updated 2026-07-25

Co-founder of the Collective Intelligence Project (CIP) and research scientist on Anthropic's Societal Impacts team. Researcher focused on participatory AI governance, collective decision-making, and AI alignment methodologies. Previously at DeepMind's Long-Term Strategy and Governance team. Co-author of 'A Vision of Democratic AI' (Digitalist Papers Vol. 1) and lead author of 'Values in the Wild'.

Saffron Huang is a researcher focused on participatory AI governance, collective decision-making, and AI alignment methodologies. She is a co-founder of the Collective Intelligence Project (CIP) and, as of 2026, a research scientist on the Societal Impacts team at Anthropic (Source: https://saffronhuang.com/). She previously worked on DeepMind's Long-Term Strategy and Governance team.

Background

Huang holds an A.B. in Applied Mathematics-Computer Science from Harvard, with minors in Government and German, completed in 2020, and is originally from New Zealand (Source: https://saffronhuang.com/). She co-founded Kernel Magazine (Source: https://saffronhuang.com/).

Before co-founding CIP, Huang worked at DeepMind on the Long-Term Strategy and Governance team, with prior work on AI governance, safety, and long-term strategy. Her own account describes her DeepMind role as a research engineer working on language models, human-AI interaction, conceptual reasoning, value alignment, and multi-agent reinforcement learning (Source: https://saffronhuang.com/). The CIP biography adds that she conducted technology governance research with organizations including the Berkman Klein Center (Source: https://www.cip.org/saffron). During this period she contributed to several DeepMind language-model papers, including "Scaling Language Models: Methods, Analysis & Insights from Training Gopher" (2021) and "Red Teaming Language Models with Language Models" (EMNLP 2022) (Source: https://saffronhuang.com/).

Huang states that she was on the founding team of the UK AI Security Institute (formerly the AI Safety Institute) and helped start its Societal Impacts team (Source: https://saffronhuang.com/).

Roles and work

At CIP, Huang co-designed Alignment Assemblies with Divya Siddarth, sortition-based deliberative processes for eliciting public input into AI governance decisions (see Alignment Assemblies and Collective Constitutional AI).

In 2023 she co-led the Collective Constitutional AI project with Anthropic. The publicly-drafted constitution produced a less-biased Claude variant that matched the capability of a researcher-drafted baseline (see Constitutional AI). The project ran a public input process involving roughly 1,000 Americans to draft a constitution for an AI system; Huang is listed as a joint co-author of the resulting paper, "Collective Constitutional AI: Aligning a Language Model with Public Input" (2024), and the work was reported on by The New York Times (Source: https://saffronhuang.com/).

As of 2026, Huang is a research scientist on Anthropic's Societal Impacts team, which she describes as studying how AI will change society to guide AI development (Source: https://saffronhuang.com/). She is the lead author of "Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions" (COLM 2025), which uses a privacy-preserving method to extract the values exhibited by Claude 3 and Claude 3.5 models across real-world conversations and releases an associated dataset and taxonomy. From an initial 700,000-conversation sample filtered to 308,210 subjective conversations, the paper derives 3,307 AI values and 2,483 human values organized into a 266/26/5-level hierarchy across Personal, Protective, Practical, Social, and Epistemic domains, and finds expression heavily concentrated — helpfulness, professionalism, transparency, clarity, and thoroughness together account for nearly 24% of occurrences. A VentureBeat report on the paper described it as an analysis of roughly 700,000 Claude conversations (Source: https://venturebeat.com/ai/anthropic-just-analyzed-700000-claude-conversations-and-found-its-ai-has-a-moral-code-of-its-own/). The taxonomy became the input to Anthropic's July 2026 follow-up, which compressed it into four value axes (Claude's values across models and languages (Anthropic Societal Impacts, July 2026)). Her Anthropic-affiliated co-authored work also includes "Clio: Privacy-Preserving Insights into Real-World AI Use" (2024) (Source: https://saffronhuang.com/). She continues to serve CIP as a technology advisor (Source: https://www.cip.org/saffron).

She is a co-author of "A Vision of Democratic AI," published with Siddarth and Audrey Tang in September 2024 as part of the Digitalist Papers Vol. 1 (see The Digitalist Papers (Stanford, Volumes 1–2)).

Writing and positions

Huang has written essays on technology and AI governance for outlets including Noema, WIRED, The New Statesman, and The Point (Source: https://saffronhuang.com/). In a Noema essay, "Here's How To Share AI's Future Wealth," she addresses concerns that AI could concentrate wealth and erode the economic value of human work, and argues for mechanisms to distribute the gains from AI more broadly (Source: https://www.noemamag.com/heres-how-to-share-ais-future-wealth/). She and co-authors have also argued, in "Generative AI and the Digital Commons" (2022), that generative models trained on publicly available data may degrade the "digital commons" they depend on and lack processes to return captured value to data producers (Source: https://saffronhuang.com/).

With Lujain Ibrahim and others, Huang co-authored "Towards Interactive Evaluations for Interaction Harms in Human-AI Systems," published by the Knight First Amendment Institute in June 2025, which proposes an evaluation paradigm for harms that develop through repeated human-AI interaction rather than from static benchmarks (Source: https://knightcolumbia.org/authors/saffron-huang).

In September 2024, Huang and Siddarth were named to the TIME 100 Most Influential People in AI list for their work founding CIP, which the magazine described as a nonprofit working to make AI development more democratic (Source: https://time.com/collections/time100-ai-2024/7012847/saffron-huang-divya-siddarth/).

Relationships