Dylan Hadfield-Menell is an associate professor at MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL), where he leads the Algorithmic Alignment Group. His research covers AI safety, alignment, cooperative inverse reinforcement learning (CIRL), reward hacking, and the technical implementation of value-aligned agents.
Background
Hadfield-Menell earned his PhD from UC Berkeley under Stuart Russell, Anca Dragan, and Pieter Abbeel. His doctoral thesis was titled The Principal-Agent Value Alignment Problem in Artificial Intelligence (Source: https://people.csail.mit.edu/dhm/). He is the son of Gillian K. Hadfield.
Roles and affiliations
Hadfield-Menell is an Associate Professor of Electrical Engineering and Computer Science (EECS) at MIT, on the faculty of Artificial Intelligence and Decision-Making within CSAIL (Source: https://people.csail.mit.edu/dhm/). He runs the Algorithmic Alignment Group, which according to his faculty page works on alignment challenges in multi-agent systems, human-AI teams, and societal oversight of machine learning, with a stated goal of enabling the safe, beneficial, and trustworthy deployment of AI in real-world settings (Source: https://people.csail.mit.edu/dhm/). Earlier institutional listings, including a 2022 Schmidt Sciences fellowship profile, described his title as Assistant Professor (Source: https://ai2050.schmidtsciences.org/fellow/dylan-hadfield-menell/); the current MIT and Cambridge Boston Alignment Initiative pages list him as Associate Professor (Source: https://people.csail.mit.edu/dhm/) (Source: https://www.cbai.ai/dylan-hadfield-menell).
He was named a 2022 Early Career Fellow under the AI2050 program of Schmidt Sciences, which placed his work under its "Alignment" hard problem; his fellowship project concerned building AI systems that manage uncertainty about rewards and adapt the support of the reward distribution in coordination with the system's ability to influence the state of the world (Source: https://ai2050.schmidtsciences.org/fellow/dylan-hadfield-menell/). He is also listed as a faculty mentor for the Cambridge Boston Alignment Initiative (Source: https://www.cbai.ai/dylan-hadfield-menell).
Research
Hadfield-Menell's earlier work includes cooperative inverse reinforcement learning (CIRL), developed with Stuart Russell, Pieter Abbeel, and Anca Dragan; research on reward hacking and goal mis-specification; and AI alignment with uncertain reward functions.
CIRL, introduced in a 2016 paper with Dragan, Abbeel, and Russell, frames the value-alignment problem as a cooperative, partial-information game between two agents, a human and a robot, both rewarded according to the human's reward function, which the robot does not initially know (Source: https://arxiv.org/abs/1606.03137). The authors argue that, in contrast to classical inverse reinforcement learning where the human is assumed to act optimally in isolation, optimal CIRL solutions produce behaviors such as active teaching, active learning, and communicative actions, and they show that computing optimal joint policies can be reduced to solving a partially observable Markov decision process (POMDP) (Source: https://arxiv.org/abs/1606.03137).
A related 2016 paper with the same co-authors, "The Off-Switch Game," analyzes the incentives an agent has to allow itself to be switched off. The authors argue that a traditional agent that takes its reward function for granted has an incentive to disable its off switch, except where the human is assumed to be perfectly rational, and that an agent will preserve its off switch only if it is uncertain about the utility of outcomes and treats the human's actions as observations about that utility; they conclude that giving machines an appropriate level of uncertainty about their objectives leads to safer designs (Source: https://arxiv.org/abs/1611.08219).
His subsequent publications span several lines of alignment and machine-learning safety research, including reinforcement learning from human feedback (RLHF), AI auditing and evaluation, interpretability, and recommender systems. He is a co-author of "Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback" (2023), "Black-Box Access is Insufficient for Rigorous AI Audits" (2024), and "Building Human Values into Recommender Systems: An Interdisciplinary Synthesis" (2024) (Source: https://people.csail.mit.edu/dhm/). Work co-authored with Gillian K. Hadfield includes "Incomplete Contracting and AI Alignment" (2019) and "Legible Normativity for AI Alignment: The Value of Silly Rules" (2019) (Source: https://people.csail.mit.edu/dhm/).
He co-authored Building AI for the Democratic Matrix (Knight Columbia, March 3 2026) with Gillian K. Hadfield and Rakshit Trivedi (Building AI for the Democratic Matrix: A Technical Research Agenda for Normative Competence and Normative Institutions (Hadfield + Trivedi + Hadfield-Menell, Knight Columbia, March 3 2026)). His contribution sets out a technical implementation pathway for normative competence, comprising sanction-detection mechanisms, attribution mechanisms, and behavioral-adjustment mechanisms in AI agents.
Relationships
- affiliated-with: Mit (CSAIL), Algorithmic Alignment Group
- family: Gillian K. Hadfield (son)
- co-author-with: Gillian K. Hadfield, Rakshit Trivedi, Stuart Russell, Anca Dragan, Pieter Abbeel (CIRL)
- author-of: Building AI for the Democratic Matrix: A Technical Research Agenda for Normative Competence and Normative Institutions (Hadfield + Trivedi + Hadfield-Menell, Knight Columbia, March 3 2026)
- supports: Normative Competence (technical-implementation perspective), AI Alignment
- related: Agent Architecture Patterns, Agentic AI, Reward Hacking
Sources
- Building AI for the Democratic Matrix: A Technical Research Agenda for Normative Competence and Normative Institutions (Hadfield + Trivedi + Hadfield-Menell, Knight Columbia, March 3 2026)
- MIT CSAIL faculty page: https://people.csail.mit.edu/dhm/
- AI2050 (Schmidt Sciences) fellow profile: https://ai2050.schmidtsciences.org/fellow/dylan-hadfield-menell/
- Cambridge Boston Alignment Initiative bio: https://www.cbai.ai/dylan-hadfield-menell
- "Cooperative Inverse Reinforcement Learning" (2016): https://arxiv.org/abs/1606.03137
- "The Off-Switch Game" (2016): https://arxiv.org/abs/1611.08219