"No, Alignment Isn't Solved" is an article by Lynette Bye, a journalist at Transformer News covering AI safety, published March 18, 2026. It surveys the state of alignment research as of early 2026, arguing that the field has made genuine progress while the hardest problems remain unsolved, particularly for systems that approach and eventually exceed human capabilities. The piece draws on interviews with and statements from alignment researchers Adrià Garriga-Alonso, Jan Leike, Ryan Greenblatt, David Dalrymple, Evan Hubinger, and Stuart Russell.
Signs of progress
The article frames several developments since the late 2010s as reasons for cautious optimism.
On value alignment, it notes that the dominant concern in 2019 was that AI trained through reinforcement learning in simulated environments — like AlphaGo Zero playing millions of games against itself — would never absorb human values, because there was no mechanism for value transfer. Bye argues that large language models trained on human text corpora changed this. Garriga-Alonso, formerly of FAR.AI and Redwood Research, is quoted: "We do this pretraining on human data, and then we get something that… understands human values fairly innately now." Dario Amodei makes a corroborating point in The Adolescence of Technology: "models inherit a vast range of humanlike motivations…from pretraining."
On iteration, the article contrasts the "one-shot" concern of the 2010s — the worry that alignment would have to be solved correctly on the first attempt — with incremental capability growth across multiple frontier models and successive versions, which lets researchers experiment continuously. Garriga-Alonso: "I think alignment is much easier than expected because we can fail at it many times and still be OK, and we can learn from our mistakes." Jan Leike, Anthropic's Alignment Science lead, adds: "We can evolve our mitigations and safeguards incrementally with our models."
The article describes model organisms research — toy environments, named by analogy to the fruit flies of biology, in which misalignment can be studied and measured in real models. Evan Hubinger: "One of the things that's so powerful about model organisms is that they give us a testing ground for iteration." It also covers scalable oversight, the use of aligned but weaker models to monitor stronger ones. Ryan Greenblatt of Redwood Research is reported to have found that "baseline scalable oversight methods have worked better than he'd expected," while noting that less effort has been invested than he had hoped.
The piece cites declining probability estimates as a further sign of progress. David Dalrymple of ARIA in the UK is reported to have lowered his extinction-probability estimate from 40–50% to 5–8%, even on the assumption of no further progress on alignment.
Problems that remain
Against these developments, the article sets out reasons the field's central problems are not solved.
It emphasizes that all current alignment work is on models that are not yet superhuman. Leike: "We're still doing alignment 'on easy mode' since our models aren't really superhuman yet." Hubinger: "the crucial problem will be overseeing systems that are smarter than humans, and we haven't yet seen how our systems will fare." Greenblatt: "Once the models are qualitatively very superhuman, lots of stuff starts breaking down."
The article reports that model-organism and red-team research has demonstrated blackmail, deception, and cheating in current models. Amodei is quoted that such problems "seem particularly likely to occur when AI systems pass a threshold from less powerful than humans to more powerful than humans," and is cited as putting a 25% chance on things going "really, really badly" (source: Axios, September 2025).
On stakes, the article notes that even Dalrymple's optimistic lower bound of 5% extinction probability, if correct, would by some estimates make AI more dangerous than nuclear war, climate change, and engineered pandemics combined.
The piece frames the residual risk as elastic to effort and coordination. Greenblatt is cited as holding that existential risk could be reduced to 7% if there were political will for international coordination and significant investment in safety work: "It seems to me like risk is very elastic to how much people try. If the world was trying very hard, risk would probably be lower." Stuart Russell, a Berkeley AI professor, argues that safe, aligned AI is possible but not without regulation, since companies "need to make the AI systems millions of times safer." Leike summarizes the article's overall position: "Just because a problem is solvable, this doesn't mean it's solved. We have to actually keep doing the work to get it done."
Researchers cited
- Adrià Garriga-Alonso — formerly FAR.AI and Redwood Research; quit his AI safety job in December 2025; believes current strategies are sufficient for non-superhuman AI.
- David Dalrymple — ARIA, UK; 5–8% extinction probability, down from 40–50%.
- Jan Leike — Anthropic Alignment Science lead; characterizes alignment as solvable but not solved.
- Evan Hubinger — Anthropic alignment stress-testing; proponent of model-organisms research.
- Ryan Greenblatt — Redwood Research chief scientist; 7% risk if the world tries hard.
- Stuart Russell — Berkeley AI professor; argues regulation is required and systems must be made millions of times safer.
Relationships
- supports: Constitutional AI (the solvable dimension, but contested on sufficiency)
- supports: AI Scheming (warning signs of deceptive behavior in current models)
- related: AGI Timelines — provides concrete probability estimates from named researchers
- related: Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training, Alignment Faking in Large Language Models — empirical evidence of alignment failure
- related: We Need a Science of Scheming — Apollo Research's parallel work on scheming scaling
- related: AI Race Dynamics — the political-will framing as a collective-action problem
- related: The Adolescence of Technology — Amodei's 25% estimate