Regulatory Markets: The Future of AI Governance is a law-review article by Gillian K. Hadfield (Johns Hopkins University and the Vector Institute) and Jack Clark (Anthropic). It proposes that governments license private regulators and require the targets of AI regulation to purchase regulatory services from them, as an addition to the AI governance toolkit rather than a replacement for existing approaches. The article was first posted to arXiv as a preprint on April 11, 2023 (arXiv:2304.04914) and revised through a fifth version dated February 3, 2026; it was published in Jurimetrics: The Journal of Law, Science and Technology, vol. 65, pp. 195–240, in Winter 2026. An earlier and shorter version of the argument appeared as Clark and Hadfield, "Regulatory Markets for AI Safety" (arXiv:2001.00078), released in December 2019, a few months after OpenAI's release of GPT-2.
The two deficits
The article's organizing claim is that AI governance in Western democracies suffers from two problems that pull in opposite directions, and that each existing approach fails at least one of them.
By a technical deficit the authors mean a lack of technical detail informing AI developers and deployers about the required operational characteristics of the systems they build and use. Saying an AI system must be "fair" does not say what fairness translates to in system performance, particularly given the multiplicity of statistical fairness measures and the impossibility of satisfying them all; the article makes the same point for "explainable" and "robust." Hadfield and Clark contrast this with conventional product-safety domains, citing Canada's toy-flammability standard — a doll or soft toy fails if a sample of its outer fabric, held at 45 degrees, ignites within one second of flame contact and the flame travels 127 millimetres in seven seconds or less — and the RTCA-published DO-178C standard that the FAA accepts as a means of compliance for aviation software. They argue that both the EU AI Act (Regulation 2024/1689) and the NIST AI Risk Management Framework 1.0 require developers to identify, assess and mitigate risks without specifying what constitutes a harm, what level of risk is acceptable, or how it is to be measured.
By a democratic deficit they mean delegating the fundamentally political task of reconciling trade-offs in AI design and deployment to politically unaccountable private actors. The article's concern is not the procedures standard-setting organizations follow, and the authors state explicitly that the problem cannot be solved by making those bodies more open: the objection is to the premise that the standards needed for AI are properly generated by private actors not accountable to the public. Standard-setting organizations, they note, originated in nineteenth- and early twentieth-century coordination on measurement and design — resistance coils, screw threads, steel rails — where the public interest was uncontested, and they cite Yates and Murphy's account of riverboat boiler standards developed by the private Franklin Institute in 1836 as a case where democratic oversight was unnecessary. AI standards, by contrast, reach core human values.
The article works the argument through two ISO documents. ISO/IEC 42001 — AI Management System requires a business to establish a responsible-AI policy and objectives such as "fairness" and "explainability," conduct impact assessments covering "human rights" and "norms, traditions, culture and values," and document objectives and log system behaviour — all relative to the policy the business itself has set, leaving the technical detail internal. ISO/IEC TR 24027, Bias in AI Systems and AI Aided Decision Making, published in November 2021, discusses sources of bias and sets out statistical metrics; the authors argue its concepts are controversial and that discrimination is addressed elsewhere in the economy through legal analysis of disparate impact and intentional discrimination, adjudicated by regulators, judges and juries who possess democratic legitimacy, not by engineers. They also record that ISO standards are proprietary, noting that they purchased the fourteen AI standards ISO had produced as of summer 2022 for approximately $2,500 under a licence forbidding copying or sharing, and that a March 2024 European Court of Justice decision (Case C-588/21 P, Public.Resources.Org and Right to Know v. Commission) held that standards referenced in European law as a means of demonstrating compliance must be made freely available.
Survey of the governance landscape
Part II reviews AI governance as of 2025. The article characterizes the preceding decade as the era of soft law and AI ethics, citing Gutierrez and Marchant's count of 634 soft law programs — defined as programs setting "substantive expectations that are not directly enforceable by government[s]" — approximately 95% of them published between 2015 and 2019, and AlgorithmWatch's inventory of 167 frameworks and principles as of April 2020. Jobin and co-authors' survey is cited for convergence around five high-level concepts: transparency, justice and fairness, non-maleficence, responsibility and accountability, and privacy. Gutierrez and Marchant predicted in 2021 that soft law would remain the dominant form of AI governance for the foreseeable future.
The authors treat the European Union as the comprehensive-legislation path, beginning with the GDPR's automated-decision provisions and continuing through the 2022 Digital Services Act and Digital Markets Act to the EU AI Act (Regulation 2024/1689). They treat the United States and United Kingdom as relying primarily on voluntary and sectoral approaches. Their conclusion is that the two paths differ less than they appear: with the exception of the DSA and DMA, all current approaches are forms of management-based or risk-based regulation in which industry identifies and manages the risks of its own products, and the most significant difference between the European and British/American approaches is that the European Union will require compliance with standard-setting-organization standards while the United States and United Kingdom will encourage and facilitate it.
The article records government moves toward direct participation in technical standard-setting: Executive Order 14110 (RESCINDED) directed NIST to develop guidance and benchmarks for evaluating and auditing AI capabilities relevant to cybersecurity and biosecurity and for developer red-teaming, and directed the Secretary of Energy to develop model evaluation tools and testbeds for nuclear, nonproliferation, biological, chemical, critical-infrastructure and energy-security capabilities. The authors note the order was rescinded in January 2025 and that the Trump administration's America's AI Action Plan, issued in August 2025, recommended NIST launch domain-specific standard-setting and support an "AI evaluations ecosystem." They cite the United Kingdom's founding of the UK AI Safety Institute (AI Security Institute) in November 2023 as grounded in the "conviction that governments have a key role to play in providing publicly accountable evaluations of AI systems," and observe that both countries stop short of mandating compliance.
A section on China argues that the framing of AI policy as a race in which the West cannot regulate because China will not "belies this caricature." It traces the 2021 recommender-algorithm regulation, which required alignment with existing internet-news law and to "uphold mainstream value[s]" while also creating user rights to disable personalized recommendations and prohibiting systems causing addiction, unsafe working conditions or price discrimination, and which created an algorithm registry; the 2023 deep-synthesis rules requiring watermarking of AI-generated content and real-identity verification of users; and the August 2023 Interim Measures for the Administration of Generative Artificial Intelligence Services. The authors observe that China faces the same technical-detail problem and has likewise turned to a standards body: the Standardization Administration of China translated the Interim Measures into concrete requirements including a keyword library of at least 10,000 words, a generated-content question bank of no fewer than 2,000 questions, and test banks of at least 500 questions each for content the model should and should not refuse, with a refusal rate of at least 95% on the former and no more than 5% on the latter. They record that the effect on system quality and cost remains unknown, and note Angela Huyue Zhang's contrary reading of China's approach as an effort to appease industry. The distinction they draw is that Chinese law requires private companies to maintain an internal Communist Party apparatus, giving the party-state direct visibility into AI companies' operations, whereas Western governments must regulate within the boundaries of private-company autonomy.
The model
Part III places the proposal within regulatory theory. The authors group "new governance" techniques into performance-based regulation (specifying results but not methods), management-based regulation (requiring firms to assess and plan for their own risks), and meta-regulation (embedding both in a system where regulators and regulated entities continually update processes and outcomes). They build on the regulatory intermediary theory of Kenneth W. Abbott, David Levi-Faur and Duncan Snidal, who define an intermediary as "any actor that acts directly or indirectly in conjunction with a regulator to affect the behavior of a target." The stated extension is that regulatory markets recruit intermediaries not merely for expertise or cost-efficiency but to attract human and financial capital into building regulatory technologies that can keep pace with AI. The model is also traced to Hadfield's 2017 book Rules for a Flat World.
The model has three actors:
| Actor | Role |
|---|---|
| Targets | Businesses and organizations building, deploying, or integrating AI. Required by government to purchase regulatory services; free to choose and switch among licensed regulators. |
| Private regulators | For-profit and nonprofit organizations that develop and sell regulatory services. Gain authority through the regulatory contract with the target plus government authorization to collect fines or impose requirements. |
| Governments | Set required outcomes (metric-based or principle-based), license and audit private regulators against those outcomes, and regulate the market itself to sustain competition and integrity. |
The authors are explicit that private regulators compete on cost and efficiency but not on the quality of their regulatory services — that is, not on the extent to which public goals are achieved — because meeting the government-set outcomes is a condition of holding a licence. This is identified as the mechanism that makes delegation legitimate and answers the democratic deficit. Regulators that fail the government's tests risk having licences suspended, conditioned or revoked, which the article says requires governments to build enough technical expertise to make that threat realistic, to ensure sufficient scale for multiple regulators (possibly by capping any one regulator's market share), and to keep switching costs low.
The article's illustrations of regulatory technology include a private regulator of self-driving cars requiring data access and using machine learning to detect accident-risk behaviours above a threshold, or developing technology to modify the target's algorithms or data sources; a banking regulator requiring differential-privacy techniques, either prescribing specific algorithms or establishing a procedure for banks to propose techniques that survive its tests; and a regulator of facial-recognition-equipped drones requiring particular cybersecurity features and algorithmic audits of accuracy across demographic groups. Matching government-set outcomes include accident-rate or congestion thresholds for autonomous vehicles, consumer credit-access thresholds for banking, and thresholds for the likelihood that drone software could be accessed by malicious users.
The central innovation the authors claim over existing new-governance models is that government shifts to establishing the goals of regulation rather than its methods, while the methods are developed by independent private regulators who are themselves regulated — rather than left to the regulated entities, as in performance-based regulation. They distinguish this from RegTech as developed in financial regulation, where the role of private vendors has been limited to digitizing and automating existing requirements; here vendors would translate government-supplied outcomes into technical requirements.
Worked example: red-teaming frontier models
The article develops one extended application: reducing the risk that frontier models are misused to create biological weapons. Governments would announce a window for licensing independent red-teaming and evaluation companies and a date by which frontier-model developers — open-weight or closed — must enter fee-based contracts with a licensed provider. Because the science is immature, the authors expect the initial licensing scheme to be an expert group assessing red-teaming companies against a general principle such as warranting "high confidence that a model is not vulnerable to state-of-the-art adversarial efforts to increase baseline capabilities among non-state actors to produce category A, B or C bioterror toxins as classified by the U.S. Centers for Disease Control," with the meaning of "high confidence," "not vulnerable" and "state-of-the-art" left to the expert group and elaborated over time.
They select red teaming because it is already in demand: Executive Order 14110 (RESCINDED) defines "AI red-teaming" as "a structured testing effort to find flaws and vulnerabilities in an AI system," directs NIST to establish red-teaming guidelines for dual-use foundation models, and requires developers of such models to report red-team results to the Secretary of Commerce; the EU AI Act (Regulation 2024/1689) likewise references adversarial testing. The article records that developers have moved from entirely internal teams toward retaining nonprofit alignment research groups including METR and Apollo Research and consulting firms such as Gryphon Scientific, which Anthropic retained in 2023 for analysis of bioweapon acceleration and which Deloitte acquired in April 2024. Under a licensing regime, the authors expect current nonprofit evaluators either to raise substantial new philanthropic funding to scale up and compete, or to incorporate as for-profit entities and seek investment capital.
Globalization argument
Part IV argues that regulatory markets address harmonization differently from treaty-style convergence. The authors take the underlying goal of harmonization to be reducing the compliance burden on companies operating at global scale, not securing agreement among governments, and cite the failure of a global effort begun in 1992 to harmonize medical-device regulation, which after twenty years had not produced a harmonized regime.
Their worked illustration posits seven private regulators of facial recognition using three technologies: regulators 1–3 audit training data for demographic representativeness, regulators 4–6 run statistical tests on audited samples of systems in operation, and regulator 7 uses human review panels, like juries, to adjudicate qualitatively. Country A licenses all seven; Country B licenses 1–6, lacking confidence in the qualitative approach; Country C licenses 4–7, lacking confidence in ex ante data controls. A provider choosing regulator 1 gains access to A and B under training-data requirements alone; a provider choosing regulator 4 reaches all three countries under statistical tests that may vary by country. Providers thus face one or a small number of regimes while reaching multiple jurisdictions, and jurisdictions keep sovereign authority over their own goals. The authors note that a country may still face de facto pressure to align — through lobbying by a provider seeking access — but say this comes without conceding sovereignty de jure, and that a less wealthy country could free-ride on the oversight efforts of larger ones. They observe that a regulatory-markets approach does not require the United States and China to agree on technical standards, since each licenses only regulators achieving its own outcomes, while noting that global standards could still emerge as a condition of market access, as WTO membership required changes of China in 2001.
A further claimed benefit is that market opportunity recruits ground-level knowledge into regulatory innovation, which the authors argue is needed because the emphasis on new AI harms has obscured the ways AI disrupts existing regulatory goals across health care, financial stability and consumer markets. They cite Sandhu, Kolt and Hadfield's Regulatory Transformation in the Age of AI (CIFAR, 2023) on that point, and note the Georgetown Emerging Technology Observatory's 2024 calculation that only 2% of AI research is focused on AI safety, against calls by Bengio and co-authors for one-third of industry and academic research funding to go to safety.
Limitations the authors state
The article devotes a section to risks, and the authors state that they do not think regulatory markets will work or be appropriate in all circumstances and that the model is "not a magic bullet for the regulatory state."
- Competition failure. If only two or three companies develop a particular type of AI, there may be insufficient scale to sustain multiple regulators. Competition may also fail through concentrated market share, high switching costs, or collusion. Antitrust law could guard against monopolization, but robust competition may require market-share limits and rules reducing switching costs.
- Capture. Placing a layer between government and industry creates the risk that private regulators, selling services to AI companies, collaborate with them to cheat on government goals. The authors argue the design partially offsets this — multiple regulators mean multiple sources of data and expertise, competitors have an incentive to expose cheating, and government regulates perhaps five or ten regulators rather than a thousand companies — but state that the model works only if governments are willing to regulate private regulators, and that it cannot fix an absence of political will.
- Underfunded oversight. The authors offer two cautionary cases: credit rating agencies in the 2008 financial crisis, which were shielded from liability for rating errors with no formal government oversight; and FAA oversight in the Boeing 737 MAX crashes, which government inspectors' reports had repeatedly found underfunded and inadequate before the accidents. Pricing regulation into the market covers some cost, but government oversight still requires funding and the model cannot eliminate the budgeting problem.
- Residual technical demand on government. Although the model is meant to address the technical deficit, overseeing private regulators still requires governments to increase in-house technical expertise; the recruitment problem is mitigated rather than eliminated.
- Political displacement. Governments may come under pressure after a high-profile accident to displace private regulators and dictate regulatory detail, which if anticipated would undermine confidence in private regulators and targets' willingness to cooperate with them.
- Complexity and opacity. New actors and processes could make regulation more complex and easier for bad actors to exploit, a concern already voiced about regulatory technology in finance.
- Feasibility. Implementation would require new licensing systems and possibly new agencies, new regulatory metrics aimed at intermediaries rather than targets, a nascent sector of regulatory intermediaries attracting investment, and legislative agreement — a bet on a new regulatory approach at a moment when, as the authors put it, the question is whether the United States will risk regulatory missteps that could diminish its lead over China.
Reception noted in the article
The authors record that in a January 2024 Wall Street Journal essay, former Google CEO Eric Schmidt — whom they describe as long a proponent of leaving AI regulation to industry alone — advocated a regulatory market in AI testing companies.
Relationships
- depends-on: Regulatory Capture in AI Policy, Risk-Based AI Regulation
- supports: Independent Verification Organizations (IVOs), Alignment Auditing
- contradicts: Standards-Industrial Complex (the article treats reliance on standard-setting organizations as generating a democratic deficit rather than resolving it)
- author-of / by: Gillian K. Hadfield, Jack Clark
- related: EU AI Act (Regulation 2024/1689), NIST AI Risk Management Framework 1.0, ISO/IEC 42001 — AI Management System, Executive Order 14110 (RESCINDED), America's AI Action Plan, METR, Apollo Research, UK AI Safety Institute (AI Security Institute), Jailbreaking and Red Teaming
Provenance
Retrieved on August 17, 2026 by direct HTTPS fetch of the canonical arXiv PDF route, extracted with pdfplumber across all 46 pages; a text layer was present throughout and no OCR was required. Authenticity was verified before the pull: the arXiv identifier resolves to the DataCite DOI 10.48550/arXiv.2304.04914, the arXiv record carries the Jurimetrics volume and page range, and the English Wikipedia article on Hadfield cites the same work at the same identifier with the same two authors. The verification record is at Wiki/_meta/queue/gap-scan/proposed-sources/regulatory-markets-hadfield-clark-2023.md.
The preprint and published dates are distinct and are recorded separately throughout: first posted April 11, 2023; fifth version February 3, 2026; published in Jurimetrics 65:195–240 in Winter 2026. Where the article dates its own survey, it describes the landscape "circa 2025."