AI Policy Wiki
Dashboard

The Emergence of Artificial Intelligence Ethics Auditing

medium confidence · updated 2026-06-06

Interview-based study of 34 AI ethics auditors across 7 countries — finds audits follow financial auditing structure but lack stakeholder involvement, success measurement, and external reporting; hyper-focused on bias, privacy, and explainability.

Source: Daniel S. Schiff (Purdue), Stephanie Kelley (Saint Mary's University), Javier Camacho Ibáñez (Universidad Europea Madrid), Big Data & Society, 2024

"The Emergence of Artificial Intelligence Ethics Auditing" is a 2024 empirical study of the AI ethics auditing field, based on 34 semi-structured interviews with AI ethics auditors across seven countries. It is one of the few empirical analyses of how AI ethics auditing works in practice rather than how it should work in theory, covering auditors' motivations, activities, and challenges. The authors describe the research as conducted in anticipation of regulatory frameworks, including the EU AI Act and state AI laws, that increasingly require or incentivize internal and external auditing of AI systems.

Findings

Structure follows financial auditing stages

The authors find that AI ethics audits follow the four stages of financial auditing (planning, fieldwork, reporting, and follow-up), borrowing legitimacy from an established professional practice. They identify three areas where the analogy to financial auditing breaks down:

  • Stakeholder involvement — financial audits require extensive stakeholder engagement, while AI ethics audits typically lack this, especially for affected communities.
  • Measurement of success — financial audits have clear quantitative metrics (whether the balance sheet is accurate), while AI ethics audits lack agreed success criteria.
  • External reporting — financial audits produce standardized public documents, while AI ethics audit reports are often internal and non-standardized.

Topic focus on technical issues

The paper reports that audits are "hyper-focused on technically oriented AI ethics principles of bias, privacy, and explainability, to the exclusion of other principles and socio-technical approaches." The authors attribute this to regulatory emphasis on technical risk management, noting that the EU AI Act and state laws focus on measurable technical properties. Typical coverage includes algorithmic bias testing, privacy assessments, and explainability or interpretability reviews. Typically excluded are broader societal impacts, labor effects, environmental harm, and power concentration — among the systemic risks documented in the GPAI taxonomy.

Challenges auditors face

The study identifies four recurring challenges: competing demands across interdisciplinary functions, since AI auditing requires expertise in machine learning, law, ethics, and domain knowledge simultaneously; firm resource and staffing constraints, with the authors noting that auditing is expensive and that firms underinvest; lack of technical and data infrastructure, with firms often unable to provide auditors sufficient access to models, training data, or decision logs; and regulatory ambiguity, where limited tractable guidance forces auditors to interpret vague requirements as best practices continue to emerge.

Constructive roles

Despite these challenges, the authors find that AI ethics auditors perform several functions: building auditing frameworks (filling a gap regulatory bodies have not filled), interpreting regulations for clients, curating practices and propagating norms across organizations, and sharing learnings with regulators as a feedback channel that influences future regulation.

Policy context

The authors argue that their findings reveal a gap between the regulatory aspiration of auditing as a meaningful check on AI systems and auditing practice in early 2024, which they characterize as largely technical, often shallow, and rarely external. They tie this to several regulatory mandates that rely on audit-like processes: the Colorado AI Act requires impact assessments for high-risk deployments; the EU AI Act requires conformity assessments for high-risk AI; the Frontier Compliance Framework includes external expert input as a component; and the "Measure" function of the NIST AI RMF is intended to be operationalized through auditing processes. The paper anticipates that the Colorado AI Act's mandated impact assessments will be technically narrow and internally focused, consistent with the audit practice it documents, and contends that if auditing remains technically narrowed and internally dominated, these mandates may not deliver the accountability they promise.

  • NIST AI RMF 1.0 — the framework's "Measure" function is primarily operationalized through auditing; this paper documents what such auditing looks like in practice.
  • Colorado AI Act — the Act's impact assessment requirements are a mandated audit-like process.
  • EU AI Act — conformity assessments for high-risk AI face the maturity challenges the paper documents.
  • Open Problems in Technical AI Governance — auditing infrastructure is identified as an open governance problem in that survey.
  • Safety Cases — external auditing is a component of safety case verification.
  • AI Safety Frameworks — auditing is the third-party verification component of safety frameworks, which are currently largely self-assessed.