Paul Christiano is an alignment researcher who serves as Head of AI Safety at the US government's frontier-model evaluation body — a position he has held since 2024, when it was the US AI Safety Institute (US AISI), and retained through the body's June 2025 transformation into the Center for AI Standards and Innovation (CAISI) (Source: en.wikipedia.org). He founded the Alignment Research Center (ARC), previously headed the alignment team at OpenAI, and has co-authored several foundational papers in RLHF and AI safety research.
Roles
| Current role | Head of AI Safety, [[nist-caisi | Center for AI Standards and Innovation (CAISI)]], NIST (since 2024, when the body was the [[us-ai-safety-institute | US AISI]]) |
| Previous | Founder, Alignment Research Center (ARC / [[metr | METR]] lineage), 2021; Head of Alignment, [[openai | OpenAI]], 2017–2021 |
| Education | BS Mathematics, MIT (2012); PhD, UC Berkeley (2017) (Source: en.wikipedia.org) | ||
| Known for | Technical contributions to [[rlhf | RLHF]] and the AI safety research agenda |
At OpenAI, Christiano led the language-model alignment team from 2017 to 2021 (Source: en.wikipedia.org), during the period leading to InstructGPT and ChatGPT. After leaving OpenAI he founded the Alignment Research Center (ARC) in 2021 as an independent alignment research organization; its evaluations arm, ARC Evals, became METR. In 2024 he was appointed head of AI safety at the US AI Safety Institute, placing an alignment researcher inside the US government's frontier-model evaluation infrastructure; NIST describes the role as designing and conducting tests of frontier AI models (Source: nist.gov). The pending appointment drew internal opposition — in March 2024, some NIST staff and scientists reportedly threatened to resign, citing concerns that his ties to effective altruism could compromise the institute's objectivity (Source: en.wikipedia.org). He has been a collaborator of Dario Amodei since their time together at OpenAI and the Concrete Problems work.
Research contributions
Christiano's work underlies alignment methods used at frontier labs. He is a co-author of *Concrete Problems in AI Safety* (Amodei, Olah, Steinhardt, Christiano, Schulman, Mané, 2016), which framed AI safety as a portfolio of concrete technical research problems. He is a lead author of "Deep Reinforcement Learning from Human Preferences" (Christiano et al., 2017), the founding paper of the RLHF technique, and a co-author of the InstructGPT paper (Ouyang et al., 2022), the founding paper of productionized RLHF, as well as a co-author of Sleeper Agents (Hubinger et al., 2024). He authored or co-authored much of the early theoretical work on iterated amplification, "AI Safety via Debate" (2018), and scalable oversight, and at ARC led work on Eliciting Latent Knowledge (ELK), a research agenda on extracting what a model internally represents even when its outputs are misleading (Source: en.wikipedia.org).
Positions on AI risk
Christiano has given public probability estimates of catastrophic outcomes from advanced AI. In 2023 he estimated a 10–20% chance of an AI takeover with many or most humans dead, and described a roughly even chance of "doom" shortly after AI systems reach human level (Source: en.wikipedia.org). These are attributed personal estimates, not consensus figures.
Relationships
- related: RLHF — key technical contributor
- related: Concrete Problems in AI Safety (co-author), InstructGPT (co-author), Sleeper Agents (co-author)
- related: OpenAI (former), METR (lineage)
- related: NIST CAISI (Center for AI Standards and Innovation) (current institutional home), US AI Safety Institute (NIST AISI) (predecessor body at appointment)
- related: Dario Amodei — longtime collaborator from OpenAI / Concrete Problems era
- supports: AI Safety Frameworks, AI Benchmarks and Evaluation