AI Policy Wiki
Dashboard

Paul Christiano

high confidence · updated 2026-07-26

Alignment researcher, founder of the Alignment Research Center (ARC), and Head of AI Safety at the US AI Safety Institute (now CAISI). Co-author of foundational RLHF and AI safety research.

Paul Christiano is an alignment researcher who serves as Head of AI Safety at the US government's frontier-model evaluation body — a position he has held since 2024, when it was the US AI Safety Institute (US AISI), and retained through the body's June 2025 transformation into the Center for AI Standards and Innovation (CAISI) (Source: en.wikipedia.org). He founded the Alignment Research Center (ARC), previously headed the alignment team at OpenAI, and has co-authored several foundational papers in RLHF and AI safety research.

Roles

Current roleHead of AI Safety, [[nist-caisiCenter for AI Standards and Innovation (CAISI)]], NIST (since 2024, when the body was the [[us-ai-safety-instituteUS AISI]])
PreviousFounder, Alignment Research Center (ARC / [[metrMETR]] lineage), 2021; Head of Alignment, [[openaiOpenAI]], 2017–2021
EducationBS Mathematics, MIT (2012); PhD, UC Berkeley (2017) (Source: en.wikipedia.org)
Known forTechnical contributions to [[rlhfRLHF]] and the AI safety research agenda

At OpenAI, Christiano led the language-model alignment team from 2017 to 2021 (Source: en.wikipedia.org), during the period leading to InstructGPT and ChatGPT. After leaving OpenAI he founded the Alignment Research Center (ARC) in 2021 as an independent alignment research organization; its evaluations arm, ARC Evals, became METR. In 2024 he was appointed head of AI safety at the US AI Safety Institute, placing an alignment researcher inside the US government's frontier-model evaluation infrastructure; NIST describes the role as designing and conducting tests of frontier AI models (Source: nist.gov). The pending appointment drew internal opposition — in March 2024, some NIST staff and scientists reportedly threatened to resign, citing concerns that his ties to effective altruism could compromise the institute's objectivity (Source: en.wikipedia.org). He has been a collaborator of Dario Amodei since their time together at OpenAI and the Concrete Problems work.

Research contributions

Christiano's work underlies alignment methods used at frontier labs. He is a co-author of *Concrete Problems in AI Safety* (Amodei, Olah, Steinhardt, Christiano, Schulman, Mané, 2016), which framed AI safety as a portfolio of concrete technical research problems. He is a lead author of "Deep Reinforcement Learning from Human Preferences" (Christiano et al., 2017), the founding paper of the RLHF technique, and a co-author of the InstructGPT paper (Ouyang et al., 2022), the founding paper of productionized RLHF, as well as a co-author of Sleeper Agents (Hubinger et al., 2024). He authored or co-authored much of the early theoretical work on iterated amplification, "AI Safety via Debate" (2018), and scalable oversight, and at ARC led work on Eliciting Latent Knowledge (ELK), a research agenda on extracting what a model internally represents even when its outputs are misleading (Source: en.wikipedia.org).

Positions on AI risk

Christiano has given public probability estimates of catastrophic outcomes from advanced AI. In 2023 he estimated a 10–20% chance of an AI takeover with many or most humans dead, and described a roughly even chance of "doom" shortly after AI systems reach human level (Source: en.wikipedia.org). These are attributed personal estimates, not consensus figures.

Relationships