Geoffrey Irving is an AI safety researcher and, since June 2026, cofounder and chief scientist of Sequent Research, a nonprofit alignment-research organization (Source: x.com; alignmentforum.org). Before cofounding Sequent he was chief scientist of the UK AI Security Institute (AISI), and earlier held research positions at DeepMind, OpenAI, and Google Brain (Source: x.com; openreview.net).
Research
Irving's work centers on technical approaches to aligning advanced AI systems, particularly scalable oversight — methods that let humans supervise models on tasks the humans cannot directly evaluate. He is a coauthor of "AI safety via debate" (Irving, Christiano, and Amodei, 2018), which proposed training agents to argue opposing sides of a question before a human judge as a mechanism for eliciting more truthful behavior than either agent could be trusted to produce alone (Source: arxiv.org). The launch announcement for Sequent credits him with an early role in establishing reinforcement learning from human feedback (RLHF) at OpenAI and DeepMind and with continued contributions to scalable oversight via debate (Source: alignmentforum.org).
At DeepMind, Irving was among the coauthors of "Ethical and social risks of harm from language models" (Weidinger et al., 2021), a taxonomy of risks posed by large language models that has been widely cited in subsequent AI-policy work (Source: scholar.google.com).
UK AI Security Institute and Sequent
As chief scientist at AISI, Irving worked on the institute's alignment research agenda. In June 2026 he and colleagues drawn from AISI's Alignment Team and the research group Timaeus launched Sequent, which argued that "alignment is not on track" for the possible near-term arrival of artificial superintelligence and that lab safety programs are unlikely on their own to deliver advance confidence that training such systems will go well. Sequent describes a federated research structure in which research directors report to Irving, and a planned initial fundraise of $100–150 million (Source: alignmentforum.org; importai.substack.com).
Relationships
- related: Sequent (Sequent Research) — cofounder and chief scientist.
- related: UK AI Safety Institute (AI Security Institute) — former chief scientist.
- related: RLHF (Reinforcement Learning from Human Feedback) — early contributor to RLHF at OpenAI and DeepMind.
- related: AI Alignment
- related: Scalable Oversight — "AI safety via debate."