AI Policy Wiki
Dashboard

Alignment Assemblies and Collective Constitutional AI

medium confidence · updated 2026-06-06

Deliberative-democracy processes — typically sortition-based — used to elicit public input into AI governance and model behavior. Core method of the Collective Intelligence Project (CIP); demonstrated across OpenAI (participatory risk prioritization, 2023), Anthropic (Collective Constitutional AI, 2023), Taiwan Ministry of Digital Affairs (moda) (Ideathons, Recursive Public, information-integrity deliberation 2024).

Alignment Assemblies are structured deliberative processes, usually sortition-based (random representative selection), designed to incorporate public values into AI governance decisions and into model behavior. They are the flagship method of the Collective Intelligence Project (CIP), founded by Divya Siddarth and Saffron Huang. The method has been applied with OpenAI (participatory risk prioritization, 2023), Anthropic (Collective Constitutional AI, 2023), and Taiwan's Ministry of Digital Affairs (moda), and it informed the alignment tuning of Taiwan's TAIDE model (The Digitalist Papers (Stanford, Volumes 1–2)).

Design

An alignment assembly is defined along four design criteria: a defined outcome, the relevant polity, the scope of discussion, and the tools and process used (The Digitalist Papers (Stanford, Volumes 1–2)). The approach draws on sortition and citizen-assembly traditions in deliberative democracy, including Fishkin's deliberative polling and the Irish Citizens' Assembly (The Digitalist Papers (Stanford, Volumes 1–2)). Deliberation has been mediated by tools including Pol.is, AllOurIdeas, and custom deliberation platforms (The Digitalist Papers (Stanford, Volumes 1–2)).

In relation to existing frameworks, the method extends Anthropic's Constitutional AI by replacing researcher-drafted principles with publicly-drafted ones (The Digitalist Papers (Stanford, Volumes 1–2)).

Documented instances

CIP and OpenAI (June 2023)

In a participatory risk-prioritization exercise, CIP and OpenAI surveyed 1,000 demographically representative Americans using the AllOurIdeas wiki-survey platform, prompting participants with "When it comes to making AI safe for the public, I want to make sure...". Concerns about public understanding and over-reliance dominated the results, contrasting with the industry focus on existential and bioweapon risks (The Digitalist Papers (Stanford, Volumes 1–2)).

CIP and Anthropic (2023): Collective Constitutional AI

Representative Americans drafted a constitution for Claude, and the publicly-drafted model was tested against the Anthropic-drafted Constitutional AI baseline. The public model was found to be less biased but equally capable, in one of the first instances where members of the public collectively directed LLM behavior. Consensus was broader than expected: more than 75% agreed AI should protect free speech, and roughly 90% agreed AI should not be racist or sexist (The Digitalist Papers (Stanford, Volumes 1–2)).

CIP and Taiwan moda (2023–2024)

CIP partnered with Taiwan's Ministry of Digital Affairs (moda) on a sequence of deliberative exercises (The Digitalist Papers (Stanford, Volumes 1–2)). An Ideathon in Taipei and Tainan addressed AI-governance and sector-transformation priorities, with emphasis on a public-sector pioneering role. Recursive Public (November 2023) was a Pol.is-mediated deliberation on AI governance. An Information Integrity Assembly (March 2024) convened hundreds of thousands of randomly-invited citizens to deliberate on platform AI labeling, provenance, and AIEC assessment criteria in the leadup to Taiwan's 2024 presidential election.

Taiwan's TAIDE model

TAIDE is an open-source LLM trained by Taiwan's National Applied Research Laboratories and built on Meta Llama 2. Its alignment-tuning phase references a constitution from the 2023 Alignment Assembly; an example principle calls for nondiscrimination on the basis of gender, religion, race, class, party, language, nationality, property, and education (The Digitalist Papers (Stanford, Volumes 1–2)).

Debates and tensions

Several questions about the method are documented (The Digitalist Papers (Stanford, Volumes 1–2)). On scale, most assemblies cap at roughly 1,000 participants, raising questions about legitimacy at nation-state scale. On legitimacy inheritance, decisions routed through corporate labs such as OpenAI and Anthropic lack government legitimacy. On manipulation risk, sortition reduces but does not eliminate lobbying and capture. Lawrence Lessig's vetocracy critique holds that assemblies can function as democratic deliberation protected from engagement and funding distortion.

Relationships