AI Policy Wiki
Dashboard

Alignment Research Center (ARC)

high confidence · updated 2026-06-06

Nonprofit alignment research organization founded by Paul Christiano after leaving OpenAI; home of the Eliciting Latent Knowledge (ELK) research agenda. ARC Evals spun off as METR.

The Alignment Research Center (ARC) is a 501(c)(3) research nonprofit based in Berkeley, California, founded in 2021 by Paul Christiano after he left OpenAI. ARC's stated mission is to align future machine-learning systems with human interests. It is best known for the Eliciting Latent Knowledge (ELK) research agenda and for early third-party dangerous-capability evaluations of frontier models, which were spun out as METR.

FieldValue
Type501(c)(3) research nonprofit
HeadquartersBerkeley, California
Founded2021
Founder[[paul-christianoPaul Christiano]]
Known forEliciting Latent Knowledge (ELK) research agenda; early frontier-model dangerous-capability evaluations (later spun out as [[metrMETR]])

Overview

ARC pursued theoretical alignment research, most prominently Eliciting Latent Knowledge, in parallel with early empirical dangerous-capability evaluations of frontier systems. Christiano founded ARC in 2021 after leaving OpenAI, where he had headed the alignment team and co-authored InstructGPT and helped define productionized RLHF.

The organization's two arms diverged over time. The evaluations arm, ARC Evals, conducted some of the first third-party dangerous-capability evaluations on frontier models in 2022 and 2023, including pre-deployment access to GPT-4 to probe autonomous-replication and resource-acquisition capabilities. In 2024 ARC Evals spun out as METR (Model Evaluation and Threat Research), a separate nonprofit operating as a principal independent evaluator for frontier labs and government. The remaining theoretical arm has continued work on ELK and related alignment-theory research since 2024, at a smaller footprint than its spinoffs.

Christiano left ARC in 2023 to lead AI safety at the US AI Safety Institute.

Eliciting Latent Knowledge (ELK)

ELK is ARC's flagship research problem, posed in Christiano, Cotra, and Xu (2021): given a model that has internal representations of facts about the world, how can humans reliably elicit that knowledge, particularly in cases where the model might have incentives to deceive? ELK is one of the canonical statements of the scalable-oversight problem and is widely cited in frontier-lab alignment research, including work at Anthropic and Google DeepMind.

The ELK challenge, a paid contest soliciting proposals, seeded a generation of alignment researchers and remains a reference framing for interpretability, honesty, and oversight research.

Relevance to AI policy

Third-party dangerous-capability evaluation as a policy concept originates in part from ARC Evals' pre-deployment work with OpenAI on GPT-4, which demonstrated that independent evaluation was operationally feasible. That model was later formalized through METR, Apollo, and the AISI network.

ELK also functions as a safety-case ingredient: Safety Cases for Frontier AI treats knowledge elicitation as a core open problem for any credible safety argument.

People

Paul Christiano (Paul Christiano) is the founder and is now at the US AISI. Ajeya Cotra and Mark Xu co-authored the ELK report. ARC's theoretical team remains small; former members populate METR, the US AISI, Anthropic, and Open Philanthropy.

Funding

ARC has been funded primarily by Open Philanthropy and other AI-safety-focused philanthropists. The evaluations arm, now METR, attracted additional funding from government and lab-adjacent sources before spinning out.

Relationships