AI Policy Wiki
Dashboard

Chris Olah

high confidence · updated 2026-06-06

Anthropic co-founder and interpretability research lead; foundational figure in mechanistic interpretability; co-founded Distill; ex-OpenAI, ex-Google Brain. At the May 2026 Vatican AI ethics conference he stated there is a real possibility AI will displace human labor at very large scale.

Chris Olah is a Canadian AI researcher, a co-founder of Anthropic, and the lead of its interpretability research program. His work from roughly 2014 onward built the field now called mechanistic interpretability, of which he is widely regarded as the founding practitioner. He previously worked at Google Brain and OpenAI, and co-founded the Distill journal.

Type: Individual (industry researcher) Affiliations: Anthropic (co-founder; interpretability research lead); formerly OpenAI, Google Brain; co-founder of Distill

Background

At Google Brain, Olah co-founded the Distill journal (2016), dedicated to clear visual explanations of machine-learning research, and produced a series of visualisations of CNN feature detectors. He moved to OpenAI in 2018 to lead the Clarity team on interpretability, then co-founded Anthropic in 2021 with the Amodei siblings and other colleagues, where he has led the interpretability research agenda since.

Research and contributions

Olah's earlier work focused on CNN feature visualisation (2015–2018), including "Feature Visualization" and "The Building Blocks of Interpretability." Through Distill, founded in 2016, he set a publication format aimed at visually and pedagogically rigorous machine-learning research. He initiated the "circuits" research program, which treats neural-network components as interpretable computational structures.

At Anthropic, Olah led the scaling of interpretability methods to frontier large language models, including sparse autoencoders and dictionary learning used to extract monosemantic features, and the "biology of an LLM" program that traces end-to-end computational pathways inside Claude 3.5 Haiku. The Anthropic interpretability team under Olah has produced the circuits-thread, features-as-directions, dictionary-learning, and the On the Biology of a Large Language Model line of work.

Selected publications and works:

  • "Feature Visualization" (Olah, Mordvintsev, Schubert, Distill, 2017)
  • "The Building Blocks of Interpretability" (Olah et al., 2018)
  • "Zoom In: An Introduction to Circuits" (Olah et al., 2020)
  • "Toy Models of Superposition" (Elhage et al., Anthropic, 2022)
  • "Scaling Monosemanticity" (Anthropic, 2024)
  • On the Biology of a Large Language Model (Anthropic, 2025)

Positions and statements

Olah's central argument is that systems whose internals cannot be inspected cannot be reliably governed or aligned, framing interpretability as a precondition for meaningful safety guarantees and for external audit. He has been less publicly positioned than Dario Amodei or Jack Clark, with commentary tending toward technical rather than political framing. His research agenda bears on what regulators and auditors consider tractable to verify, a consideration relevant to AI Safety Cases and Frameworks and evaluation-based regulation.

Vatican encyclical launch and AI ethics conference

Olah co-presented Pope Leo XIV's first encyclical Magnifica humanitas at the Vatican on May 25, 2026, a Vatican move that paired the Anthropic interpretability lead, who is not Catholic, alongside theologians and Vatican officials before an audience of cardinals, computer scientists, journalists, and diplomats, including US Ambassador to the Holy See Brian Burch. On launch day Olah described an incentive-conflict problem at frontier-AI companies:

"We need moral voices that the incentives cannot bend."

"[AI companies] need moral guidance to avoid being swayed by a set of incentives and constraints that can sometimes conflict with doing the right thing."

"Today is just the beginning — the start of a long collaboration between those of us who are building this and those who can see what we, from the inside, cannot."

No prior Anthropic principal had explicitly invoked an external moral authority as a check on lab incentives; the RSP and the Amodei and Clark public-position framings treat institutional self-restraint as the principal lever, and Olah's "moral voices that the incentives cannot bend" formulation cedes that internal frameworks are insufficient. Reception was contested. Paolo Carozza (Notre Dame, chair of the Meta Oversight Board) said that "frontier AI companies would love to co-opt religious communities to bring an ethical imprimatur to their work," while Rev. Andrea Ciucci (Pontifical Academy for Life) argued that "those who create the technology also have to bear responsibility — so we need to talk to them." The encyclical text reportedly shows Anthropic linguistic influence, describing AI software as something programmers "scaffold" rather than directly code, mirroring Anthropic's own "AI grows" framing. (Source: nytimes.com; cruxnow.com)

Speaking at the Vatican AI ethics conference connected to the Magnifica humanitas launch on May 27, 2026, Olah stated that "there is a real possibility that AI will displace human labor at very large scale" — harder language than prior Anthropic principals' public statements, and the first appearance on the record of "very large scale" labor-displacement framing from a named Anthropic principal, against the earlier Amodei "50% of entry-level white-collar jobs" framing, which has since been walked back into a Jevons-paradox-multiplier reframe. The remark came the same week Sam Altman said he was "pretty wrong" about AI's near-term white-collar impact and "delighted to be wrong," placing Anthropic and OpenAI on opposite sides of the doom-versus-hype axis they had previously occupied, as both companies headed into reported roughly $1T IPO talks; see AI Labor Disruption for the consolidated framing. (Sources: anthropic.com; axios.com)

A May 26, 2026 Wired Backchannel piece, "Why the Vatican Invited Anthropic to the Pope's AI Encyclical Presentation," carried Olah's most candid published acknowledgment of the lab-incentive problem to date, that even safety-focused labs remain "immersed in economic, geopolitical, and competitive incentives that can conflict with doing the right thing." It was the first time an Anthropic principal named "competitive incentives that can conflict with doing the right thing" in the same context as the RSP framework the lab points to as its institutional self-restraint mechanism. The piece paired this with the Pope Leo XIV line that "technological power takes on a new face, one that is predominantly private" and the encyclical's invocation of a "digital Babylon," statements relevant to the concentrated-power framing discussed under AI Power Concentration and AI Labor Disruption. (Source: wired.com)

Relationships