AI Policy Wiki
Dashboard

AI Safety vs. AI Ethics Divide

contested confidence · updated 2026-06-06

The institutional and intellectual divide between the AI safety (x-risk) field and the AI ethics (near-term harms) field — origins, TESCREAL critique, post-2023 partial convergence, and the tensions that persist.

For roughly a decade (≈2015–2023), AI research concerned with harms was split between two intellectual communities: the AI safety field, focused on catastrophic and existential risk from advanced AI, and the AI ethics field, focused on present-day harms from deployed systems such as discrimination, labor impacts, surveillance, and environmental cost. The two fields had different foundational texts, funding sources, institutional homes, demographic compositions, and theories of change. From 2023 to 2025 partial convergence emerged, with both communities addressing many of the same policy questions, but underlying tensions persist in ways that shape debates over the EU AI Act, the California bills, and the broader regulatory landscape.

This characterization is itself contested: both communities contain members who would reject the framing. The description below is intended to describe the debate rather than adjudicate it, which is reflected in the page's contested confidence rating.

Definition

"AI safety" (sometimes "long-term AI safety" or "AI x-safety") is the field studying how to prevent AI systems from causing catastrophic or existential harm, with particular attention to loss-of-control scenarios, misalignment, and dangerous capability thresholds (CBRN, autonomous replication, advanced deception).

"AI ethics" (sometimes "responsible AI" or "FAccT — Fairness, Accountability, Transparency") is the field studying how to prevent AI systems from causing harm in the near term, with particular attention to algorithmic discrimination, surveillance, labor exploitation, data-extraction harms, and disparate impacts on marginalized groups.

The distinction is institutional and conceptual rather than strictly about time horizon; near-term risks appear in both fields' work.

Intellectual origins and figures

AI ethics grew out of STS (science and technology studies), feminist technology studies, critical race scholarship, and computing-ethics traditions. Foundational texts include Ruha Benjamin's Race After Technology, Safiya Umoja Noble's Algorithms of Oppression, Cathy O'Neil's Weapons of Math Destruction, and the Gender Shades project (Buolamwini and Gebru). The ACM FAccT conference is the field's flagship venue. Figures associated with the field include Timnit Gebru, Joy Buolamwini, Meredith Whittaker, Kate Crawford, Emily Bender, Safiya Noble, and Ruha Benjamin; associated institutions include the AI Now Institute, the Distributed AI Research Institute (DAIR), and the Algorithmic Justice League.

AI safety grew out of effective altruism, rationalist communities, and certain strands of analytic philosophy. Foundational texts include Yudkowsky's sequences, Bostrom's Superintelligence, Russell's Human Compatible, and Christiano's iterated distillation and amplification papers. Venues include NeurIPS safety workshops, AI Safety Camp, MIRI, CHAI, and the Alignment Forum. Figures associated with the field include Eliezer Yudkowsky, Nick Bostrom, Stuart Russell, Paul Christiano, Dan Hendrycks, and Yoshua Bengio; associated institutions include MIRI, ARC, the Center for AI Safety (CAIS), Anthropic, Redwood Research, and the Center for the Governance of AI (GovAI). Several of these figures lack dedicated pages (see Missing Coverage: Companies, People, Models, Legislation, Concepts (revised)).

Theory of harm

The ethics field focuses on harms from systems being deployed now: discrimination in hiring, biased medical AI, surveillance overreach, labor displacement, and environmental cost. Its frame is typically structural and political, treating harms as products of existing power relations amplified by technology.

The safety field focuses on harms from systems more capable than current ones: catastrophic misuse, loss of control, deceptive alignment, and racing dynamics producing unsafe frontier capabilities. Its frame is typically technical and policy-oriented, treating harms as emergent properties of capable optimization.

Theory of change

The ethics field favors regulation, labor organizing, civil-rights enforcement, participatory design, and structural reform. It is skeptical of both voluntary industry commitments and fixes that leave underlying power relations intact.

The safety field favors technical alignment research, industry commitments (RSPs, preparedness frameworks), narrow targeted regulation such as compute thresholds, international coordination through summits and treaties, and evaluation infrastructure. It is more open to industry self-regulation when properly structured, and more willing to accept regulatory approaches that target only frontier systems.

Funding and institutions

The ethics field is funded predominantly by foundations (Ford, MacArthur, Mozilla, Open Society), university grants, and some corporate CSR. Its institutional homes are universities, civil-society organizations, and dedicated ethics institutes.

The safety field is funded predominantly by Open Philanthropy, the FTX Future Fund (prior to its collapse), Longview, the Survival and Flourishing Fund, and AI-company-linked sources. Its institutional homes include AI companies themselves (DeepMind safety teams, Anthropic, OpenAI safety teams), dedicated safety institutes such as MIRI and ARC, and a cluster of EA-aligned research groups. A recurring ethics-side critique points to this funding asymmetry: safety research is often funded by or adjacent to the firms whose products it is supposed to evaluate.

On demographics, observers note, contested and informally, that the safety community has been disproportionately drawn from certain EA, rationalist, and analytic-philosophy backgrounds, while the ethics community has been more demographically diverse, including more women, more Black researchers, and more non-US voices. Whether this is causal of the frame differences or coincident with them is contested.

Formative events

The Stochastic Parrots paper (Bender, Gebru, McMillan-Major, Mitchell, 2020/2021) argued that large language models posed serious ethics harms — environmental cost, bias amplification, and exclusion of marginalized voices — and should be approached with caution. Gebru's departure from Google following an internal dispute about the paper became a foundational event for the AI ethics community and a reference point for critiques of industry-captured safety work.

The CAIS Statement on AI Risk (May 2023) stated that "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war." It was signed by the three most cited AI researchers (Bengio, Hinton, and subsequently Sutskever) along with hundreds of industry and academic figures, marking a visible shift of mainstream AI leadership toward the safety community's framing.

The TESCREAL critique

TESCREAL, coined by Émile Torres and Timnit Gebru, is an acronym for Transhumanism, Extropianism, Singularitarianism, Cosmism, Rationalism, Effective Altruism, and Longtermism. The TESCREAL critique argues that these ideologies share common roots, common funders, and common preoccupations with far-future hypothetical harms, and that they distort AI governance by directing attention and resources away from present harms. It is the most developed ethics-side critique of the safety community.

Proponents of TESCREAL-aligned work respond that the acronym collapses distinct positions that serious proponents do not endorse — most alignment researchers are not committed longtermists in Bostrom's sense — and that the framing is ad hominem rather than substantive.

Post-2023 convergence

Following the 2023 shift (the CAIS Statement, the Bengio shift, the Biden EO, the Bletchley Summit, and the FLI Pause Letter), substantive convergence emerged across several dimensions. Both communities now address catastrophic misuse risk (CBRN, cyber) and structural harms (discrimination, surveillance) in their policy engagement. The EU AI Act, the UK AI Safety Institute, and California SB 53 each contain elements responsive to both framings. Several ethics-side figures, including Gary Marcus and Melanie Mitchell, now engage with frontier-risk questions, while several safety-side figures, including Hendrycks and Bengio, explicitly incorporate near-term harms in their framing. The national-security frame has absorbed elements of both, with loss-of-control and misuse concerns appearing alongside discrimination and civil-liberties considerations.

Persisting tensions

The EU AI Act is structured primarily around near-term deployment risks such as discrimination, surveillance, and conformity assessment; its frontier-model rules (GPAI duties) were a late addition and are themselves contested by both camps. California SB 1047 was opposed both by a vocal ethics contingent, which viewed it as irrelevant to present harms, and by industry, while being supported by much of the safety community; it was vetoed by Newsom and replaced by SB 53. A White House communications posture in 2025–2026 recast some safety-community advocacy as "radical activists" engaged in activist interference, reinforcing the post-2023 political reframing of both fields by Republican-aligned policy-making (Source: washingtonpost.com). The Frontier Model Forum is composed explicitly of frontier-lab safety-oriented participants, without ethics-community representation. Industry safety funding continues to exceed ethics funding.

Relationships