"Core Views on AI Safety: When, Why, What, and How" is a foundational essay published by Anthropic (institutional authorship) around March 2023 at anthropic.com/news/core-views-on-ai-safety. It sets out the company's stated worldview: why it believes transformative AI is likely within the coming decade, why it treats safety as a central priority, and how it approaches safety research empirically and as a portfolio. The essay reads as both a company position statement and a research rationale, and articulates the reasoning Anthropic gives for its own existence.
Summary of argument
The essay advances three connected claims: that rapid AI progress is likely, that the resulting safety risks are real, and that the right response is a portfolio of research bets across multiple scenarios rather than a single strategy.
Why rapid AI progress is likely
Anthropic grounds its expectation of rapid progress in four observations. It cites scaling laws — predictable capability improvements that follow from increased compute (Kaplan et al. 2020, Scaling Laws for Neural Language Models) — and notes the trend had held across multiple orders of magnitude. It argues that previously assumed barriers, including multimodality and logical reasoning, had been overcome. It observes that AI training compute was growing roughly 10× per year as of writing, with the total budget still far below "big science" projects, leaving room for further growth. Finally, it points to feedback loops within development itself: code models make AI researchers more productive, and Constitutional AI reduces dependence on human feedback, both of which accelerate development from within.
Why safety risks are real
The essay identifies two categories of risk. The first is the technical alignment problem: building safe, steerable systems becomes harder as those systems approach or exceed human-level capability. It draws an analogy to chess, where a novice cannot evaluate a grandmaster's moves, and argues the same evaluative asymmetry applies once AI exceeds human expertise. The second is structural and disruption risk: rapid progress is expected to be destabilizing for employment, macroeconomics, and international power structures, and competitive races between nations or corporations could lead to deployment of untrustworthy systems.
The portfolio approach to safety
Anthropic states that it refuses to bet on a single scenario and instead divides the future into three:
- Optimistic — safety challenges are easy; RLHF and Constitutional AI are sufficient, and focus shifts to near-term harms and structural risks.
- Intermediate — catastrophic risks are possible but intensive safety research can address them, with mechanistic interpretability potentially the decisive tool.
- Pessimistic — safety is essentially unsolvable at high capability levels, and the goal becomes building compelling evidence to halt development.
The portfolio approach aims to make meaningful progress in intermediate scenarios (where it claims the highest impact), to raise the alarm in pessimistic scenarios, and to contribute in optimistic ones.
Three types of research
The essay distinguishes three categories of work at Anthropic. Capabilities research is needed to have frontier models to study and is generally not published. Alignment capabilities are new algorithms for making AI more helpful, honest, and harmless — RLHF, Constitutional AI, automated red-teaming, and debate — and carry pragmatic value. Alignment science evaluates whether alignment capabilities actually work and how they scale; mechanistic interpretability is described as the flagship, functioning as a "red team" against alignment capabilities.
Research directions named
As of writing, the essay lists its active safety research directions as mechanistic interpretability, scalable oversight, process-oriented learning (rewarding the process rather than the outcome), understanding generalization, testing for dangerous failure modes, and societal impact evaluations.
Standing and later context
The essay predates most of the empirical alignment results published since, including Alignment Faking in Large Language Models, Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training, and Frontier Models are Capable of In-context Scheming. Its optimistic/intermediate/pessimistic framework remains a usable orienting tool. It also articulates the argument that frontier safety research requires frontier models, which is the central justification Anthropic gives for building powerful models while claiming a safety motivation.
Tensions
The essay acknowledges that doing safety research on frontier models risks accelerating dangerous capabilities. To address this it promises "externally legible commitments" and external evaluation, later formalized as the RSP.
The three-scenario framework is agnostic about which scenario holds. By 2026, evidence from Alignment Faking in Large Language Models, Agentic Misalignment: How LLMs Could Be Insider Threats, and Frontier Models are Capable of In-context Scheming points toward at least the intermediate scenario.
Relationships
- supports: AI Safety Cases and Frameworks — foundational essay for Anthropic's safety framework philosophy
- supports: Mechanistic Interpretability — explains why mechanistic interpretability is Anthropic's primary bet
- supports: Constitutional AI — Constitutional AI is identified as a promising alignment-capabilities technique
- related: AI Alignment — the motivating argument for the broader alignment research agenda
- related: Anthropic — the institutional rationale for Anthropic's existence
- related: Anthropic's Responsible Scaling Policy (Version 3.1) — the "externally legible commitment" promised in this essay, later delivered
- related: Scaling Laws for Neural Language Models — the empirical foundation cited for the rapid-progress argument