Roman Yampolskiy is a computer scientist and professor at the University of Louisville who works on AI safety and security. He is known for research on the theoretical limits of AI controllability, arguing through formal and philosophical analysis that the AI alignment problem may be fundamentally unsolvable.
Background
Yampolskiy is a professor at the University of Louisville, where his work centers on AI safety and security and, in particular, the theoretical limits of AI controllability.
Positions and statements
Yampolskiy argues that no technical alignment solution may be sufficient to guarantee safe superintelligent AI, a position he frames as more pessimistic than most others in the safety community. His arguments draw on computability theory and undecidability results, contending that verifying an AI system's alignment is at least as hard as solving the halting problem.
Yampolskiy holds that if alignment is formally unsolvable, the implication is a stronger case for strict capability controls, moratoriums, or containment-based approaches rather than reliance on alignment research alone. His arguments connect to concerns about Deceptive Alignment and Recursive Self-Improvement (RSI), and his threat model includes uncontrollable recursive self-improvement.
Relationships
- related: AI Alignment — argues alignment may be formally unsolvable
- related: Recursive Self-Improvement (RSI) — his threat model includes uncontrollable self-improvement
- related: Deceptive Alignment — relevant to his arguments about verification impossibility