Safetywashing refers to the misrepresentation of advances in general AI capabilities as advances in AI safety. The term was introduced by Richard Ren and co-authors in the 2024 paper "Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?" and has since been used more broadly to describe organizational compliance, governance processes, and public messaging presented as safety while not reducing underlying risk. It is analogous in construction to "greenwashing."
Origin in benchmark analysis
The Ren et al. paper conducts a meta-analysis of AI safety benchmarks, measuring how strongly performance on each correlates with upstream general capabilities such as knowledge and reasoning, and with training compute, across dozens of models. The authors find that many widely used safety benchmarks correlate highly with general capabilities, meaning that a more capable model scores better on them largely by being more capable rather than by being safer. Where this holds, capability improvements can be reported as safety improvements — the phenomenon the authors name "safetywashing." On that basis they argue that safety research should prioritize metrics that are empirically separable from, and not highly correlated with, generic capability gains, and they propose defining AI safety in a machine-learning context as a set of research goals distinguishable from capability advancement (Source: https://arxiv.org/abs/2407.21792). The paper was published at NeurIPS 2024; its authors are associated with the Center for AI Safety (Source: https://github.com/centerforaisafety/safetywashing).
Broader usage in governance debates
Beyond the benchmark-correlation result, the term is used in policy and governance discussion to describe situations where formal safety activity functions as a substitute for risk reduction rather than a means to it — for example, where documentation, audits, or voluntary commitments signal diligence without changing system behavior. In this sense it is treated as adjacent to Regulatory Managerialism, the critique that risk-management proceduralism can displace substantive safety, and it appears in sociotechnical-governance discussion alongside that critique (Source: https://arxiv.org/abs/2407.21792). Used this way, the concept overlaps with debates over whether benchmark scores, red-teaming reports, and safety frameworks reliably track real-world risk.
Relationships
- related: AI Safety Cases and Frameworks — frameworks whose safety claims the critique scrutinizes.
- related: Regulatory Managerialism — adjacent critique of compliance-as-substitute-for-safety.
- related: Sociotechnical AI Risk Governance — sociotechnical framing in which the term is cited.
- depends-on: NIST AI Risk Management Framework 1.0 — risk-management frameworks whose evaluative claims the critique scrutinizes.
Provenance note
Page created 2026-06-15 (gap-scan) from the arXiv abstract and project page of the originating paper. It resolves a dangling [[concepts/safetywashing]] reference relied on by Sociotechnical AI Risk Governance, Regulatory Managerialism, and the Knight Columbia risk-governance source page.