AI Policy Wiki
Dashboard

Independent Verification Organizations (IVOs)

medium confidence · updated 2026-08-17

Private third-party bodies, licensed or accredited by a government, that certify AI developers against safety outcomes the state sets rather than rules the state writes. The institutional layer common to the 'regulatory markets' and 'private governance' families of proposals, and the mechanism adopted by the FRONTIER Act, Connecticut's omnibus AI law, and California SB 813. Its recurring criticism is the developer-pays conflict of interest that discredited the credit-rating agencies before 2008.

An independent verification organization (IVO) is a private body, licensed or accredited by a government, that audits or certifies an AI developer against safety outcomes the government specifies. The arrangement divides labor between the state and the market: the state sets the outcomes it wants and decides who may verify them, while competing private verifiers develop the technical standards and perform the assessments. It is the institutional layer shared by the two main US proposals for private AI governance, and it is the mechanism written into the bipartisan FRONTIER Act, into Connecticut's 2026 omnibus AI law, and into successive versions of California SB 813.

The model's stated attraction is that legislators and agencies understand frontier systems less well than the laboratories building them, and that rules fixed in advance cannot keep pace with the technology, whereas private verifiers sit closer to the technology and are disciplined by competition (Don't Let AI Developers Hire Their Own Referees (Weil, July 2026)). Its recurring criticism concerns who pays: under most versions the developer selects and pays its own certifier.

Much of the survey below rests on Gabriel Weil's July 2026 commentary in AI Frontiers, which both catalogues the live instruments and argues against the model; his account is treated as a position throughout, and where it diverges from the primary record on a particular instrument both readings are recorded rather than reconciled.

Origin: regulatory markets and private governance

Two named framings converge on the IVO mechanism.

Regulatory markets. Gillian Hadfield developed the idea under this label, with Jack Clark: AI developers must pay for oversight from private regulators that governments license and hold accountable for meeting safety standards (Regulatory Markets: The Future of AI Governance). The proposal was first posted as an arXiv preprint in April 2023 and published in Jurimetrics in 2026, with a shorter precursor released in December 2019. Its stated motivation is a pair of problems the authors name: a technical deficit, the absence of technical detail telling developers what operational characteristics are required of the systems they build, which they argue neither the EU AI Act (Regulation 2024/1689) nor the NIST AI Risk Management Framework 1.0 supplies; and a democratic deficit, the delegation of value-laden trade-offs to politically unaccountable standard-setting bodies. In their model government sets outcomes and licenses regulators against them, while private regulators compete on cost and efficiency but not on the degree to which public goals are met, since meeting the government's outcomes is a licensing condition (Regulatory Markets: The Future of AI Governance).

Hadfield and Clark's version is distinctive in two respects that later US instruments do not carry over. First, it is framed as a global market: private regulators seek licences from each jurisdiction in which they operate, so that a developer complying with one regulator can reach several countries while each country retains authority over its own outcome requirements — an approach to harmonization aimed at reducing cross-jurisdictional compliance burden rather than securing agreement between governments. Second, its purpose is not only to certify but to attract investment into building regulatory technology, on the argument that the public sector will not build such tools itself (Regulatory Markets: The Future of AI Governance). The authors state the developer-pays and capture risks directly, offering the pre-2008 credit rating agencies and FAA oversight of the Boeing 737 MAX as cautionary cases, and say the model works only where governments are willing to fund and exercise oversight of the private regulators (Regulatory Markets: The Future of AI Governance).

The framing recurs in Risk-Based AI Regulation, where reporting on Dario Amodei's June 2026 FAA-analogy proposal described third-party evaluation as performed either by an FAA-style agency or by government-authorized private evaluators — described in that reporting as a regulatory-markets approach (Policy on the AI Exponential (Dario Amodei, June 2026)).

Private governance. Dean Ball has argued for a version he calls private governance, likening the arrangement to bank supervision (Source: hyperdimensional.co). Ball's recurring position favors narrow, evaluation-led federal governance: a regime led by CAISI conducting narrow-domain evaluations, backed by state-authorized IVOs. In a May 4, 2026 Hyperdimensional essay he set out a two-layer alternative to ad-hoc federal pre-deployment review — CAISI-led narrow-domain cyber-vulnerability evaluations as the technical layer, state-authorized IVOs as the institutional layer, with their findings serving as legitimacy proxies for federal procurement and deployment (Source: hyperdimensional.co). Ball describes the design as adjacent to financial-audit accreditation regimes. The nonprofit Fathom has converted the private-governance concept into model legislation (Source: fathom.org).

The two framings differ in emphasis rather than mechanism: Hadfield's stresses the competitive market in regulatory services, Ball's the division between a federal technical layer and a state licensing layer. Both leave the verifier private and the licensor public.

Instruments adopting the model

InstrumentVerifier designLicensor / accreditorStatus
FRONTIER Act (Obernolte–Trahan, House, July 2026)Largest frontier developers must retain licensed IVOs for regular adequacy audits of their risk-management efforts, reporting to federal overseersCommerce Department; framework administered by CAISI with a $100 million annual budgetIntroduced July 2026, 74 pages (Frontier Act / Great American AI Act (Obernolte–Trahan))
Connecticut omnibus AI law (2026)Multiyear pilot under which the state may approve up to five IVOs; certification "would help companies in court without entirely shielding them from liability"State consumer-protection departmentEnacted spring 2026 (Connecticut SB 5 — Broad AI law (frontier reporting + ADMT + AI companions + sandbox))
California SB 813Criteria for determining whether a registered AI auditor qualifies as an independent verification organization; paired with AB 1405's third-party AI-auditor registryCalifornia AI Standards and Safety Commission, under the Business and Consumer Services Agency, in the July 2, 2026 amended textContested — see below (California SB 813 (AI Standards and Safety Commission))
VirginiaState commission directed to study the IVO model for AI regulationStudy directive (Source: lis.blob.core.windows.net)
ARI federal blueprint (August 2026)Government conducts all assurance initially; accredited IVOs may augment compliance assurance only, in a three-phase shift, and only in domains the regulator certifies have capacity; adequacy assurance never delegatedFederal regulator alone, by domain; supervision modeled on the PCAOBAdvocacy proposal, not introduced (Responsible Innovation at the Frontier (ARI, August 2026))

The FRONTIER Act is the most developed federal instrument. Its predecessor, the June 2026 Great American AI Act discussion draft, already required frontier developers to retain licensed IVOs for regular adequacy audits and carried $1 million civil penalties for misrepresenting safety practices or failing to report critical incidents; the introduced bill retained the licensed-IVO third-party-audit model while narrowing the draft's three-year preemption of state AI-development law to state transparency, audit, and incident-reporting laws (Frontier Act / Great American AI Act (Obernolte–Trahan)). Its disclosure regime is auditor-facing rather than public, modeled on California SB 53 and New York's RAISE Act.

Two descriptions of the state instruments do not currently reconcile, and both are recorded rather than resolved:

  • California SB 813. Weil's July 29, 2026 account states that SB 813, backed by Fathom, "would have let developers earn a shield from tort liability if they met standards set by a private organization accredited by the state attorney general," and that "it failed this session" (Don't Let AI Developers Hire Their Own Referees (Weil, July 2026)). The California SB 813 (AI Standards and Safety Commission) page records a different structure and a different status as of July 11, 2026: a state commission developing voluntary standards, IVO qualification criteria set by that commission rather than the attorney general, no tort shield in the July 2, 2026 amended text, and the bill advancing 11–2 out of the Assembly Privacy and Consumer Protection Committee on July 1 and re-referred to Appropriations (Source: calmatters.digitaldemocracy.org). The accounts may both hold if Weil is describing the pre-amendment Fathom-backed version and a later event ended the bill between July 11 and July 29; neither reading is confirmable from the primary record held here.
  • Virginia. Fathom chief executive Andrew Freedman stated in a July 1, 2026 Transformer op-ed that Connecticut and Virginia "have already passed independent-verification bills" (NIST CAISI (Center for AI Standards and Innovation); Source: transformernews.ai), while Weil describes Virginia as having directed a commission to study the model. Virginia's enacted-or-not status for IVOs is therefore unsettled here. The Virginia instrument in question is not Virginia HB 2094 (High-Risk AI Developer and Deployer Act, vetoed), the Colorado-style high-risk bill vetoed in March 2025.

The developer-pays criticism

The most developed objection is that the party being certified chooses and pays the certifier. Weil argues this reproduces the conflict of interest that discredited the credit-rating agencies after the 2008 financial crisis: when issuers shopped for the agency that would bless their securities, competing agencies were under pressure to be lenient. An IVO dependent on the developers it clears for repeat business has an incentive to grade gently, and a developer shopping among IVOs will find the one that does — so that "competition, the feature that is meant to make the IVO model effective, instead drives it toward laxity." Where a passing grade also carries a liability shield, he argues the problem compounds, because a shield "swaps the broad incentive to cut risk by any cost-effective means for a narrow incentive to do only what earns the shield" (Don't Let AI Developers Hire Their Own Referees (Weil, July 2026)). Weil credits the credit-rating parallel to earlier critics (Source: transformernews.ai).

The August 2026 ARI blueprint is the most direct attempt to design around the developer-pays objection while retaining a private-verifier layer, and it does so by narrowing what verifiers decide and by restructuring how they are paid. It separates compliance assurance, a technical determination against a fixed published standard that may be delegated, from adequacy assurance, a determination bearing on future rulemaking that is never delegated regardless of how mature the verifier market becomes. Delegation of compliance work proceeds only through a three-phase shift gated on the regulator certifying, against fixed criteria, that sufficient accredited capacity exists in a given risk domain, and in every domain, cycle and phase the government conducts a share of examinations — a capability floor the authors justify as preserving the regulator's ability to reproduce and benchmark verifier work. Its market safeguards address the shopping problem directly: examination fees paid by developers pass through a regulator-administered fund rather than to the verifier, an IVO's revenue from any single developer is capped as a share calibrated to market size, rotation of both IVOs and government examination teams is mandated, revolving-door restrictions apply between examiners and the developers they examine, and any marked divergence between two assessments draws scrutiny that may reach the IVO's accreditation. On liability, the blueprint inverts the usual direction: protection may be extended while the market is forming and while the government examines the same domain, but diminishes once IVO examinations become load-bearing, on the argument that "there could be nothing worse for frontier AI safety than a robust IVO market that enjoys immunity from its own negligent conduct." For developers, no examination outcome creates immunity, safe harbor, or a presumption against liability — which places the proposal alongside Weil's objection to shield-carrying designs rather than against it (Responsible Innovation at the Frontier (ARI, August 2026)).

A second objection targets the licensing layer rather than the verifier. Proponents answer the incentive problem by having a government body license IVOs and revoke the licenses of those that grade too easily. Weil argues this reintroduces the problem the model was meant to solve, since policing a market of verifiers requires a public body with the expertise to second-guess technical judgments and the independence to withstand pressure from large firms — and "if the government could reliably field such a body, much of the reason to outsource verification at all would fall away."

Proponents' stock counter-example is Underwriters Laboratories, whose mark has functioned as a trusted private safety certification for over a century. Weil reads its history against the model: UL began in 1894 as the Underwriters' Electrical Bureau, backed by fire-insurance underwriters who needed honest assessments of electricity because their own capital was exposed when buildings burned (Source: ul.org). On his account the arrangement survives the manufacturer-pays structure today only because consumers of a product are usually also the people it might harm, so market demand for safety is real; whereas much frontier-AI risk falls on nonconsenting third parties — his example is a model conducting a cyberattack against a company that is neither its developer nor its user — leaving customer demand too weak a signal.

The insurance alternative

Weil's proposed substitute keeps the private-verifier structure but changes who holds the pen. Under a mandatory liability-insurance requirement, the insurer occupies the verifier's role with its own capital exposed: if it underprices it pays in claims, if it overprices it loses the client, and because its capital stays exposed for the life of the policy it keeps monitoring the developer and can condition renewal on fixes. He offers the Insurance Institute for Highway Safety as the comparison — insurers fund IIHS crash-prevention ratings because they pay auto-liability claims, giving the certifier a principal with money at stake, which he says the IVO model lacks (Source: iihs.org; Don't Let AI Developers Hire Their Own Referees (Weil, July 2026)).

He argues the design loses nothing in expertise, since an insurer can hire the same specialists an IVO would or contract the assessment to a firm answering to a principal that loses money if the assessment is wrong; that it extends the reach of liability, because the largest AI harms would exceed the liquidation value of the developers and judgments above that value deter nothing; and that unlike safe-harbor versions of the IVO model it confers no immunity, leaving the developer liable and insured rather than shielded. Government would still set outcomes and license verifiers, but the tasks reduce to setting minimum coverage and applying standard solvency and conduct rules — neither requiring the state to judge which models are safe.

The proposal's stated limit is the tail. Premiums are an honest signal only where the insurer's capital is at stake, and for extreme low-probability catastrophes the insurer is itself judgment-proof, so those risks do not reach the premium. Weil therefore confines the requirement to what he calls the "insurable layer," and names three other tools for the rest: mandatory disclosure of insurer risk assessments with liability for false ones, shared residual liability across frontier developers on the model of the Price-Anderson Act's retrospective assessments, and a public backstop on the model of the Terrorism Risk Insurance Act, whose cap limiting payouts to $100 billion per year he describes as a design choice Congress can reset. His summary position: "The insurer is the right verifier for the insurable layer. It is not, by itself, the answer for the most extreme risks."

The argument runs against the third-party-audit architecture that federal and state instruments have converged on. It also cuts against a separate institutional argument from the same year: Freedman's reading of the Supreme Court's June 29, 2026 Slaughter decision, which held the president may remove independent-agency heads at will, was that the ruling strengthens the case for accredited IVOs reporting to CAISI, since freestanding independent agencies can no longer be insulated from at-will removal (Trump v. Slaughter).

Relation to government pre-release review

IVOs are frequently proposed as the legitimate alternative to direct government vetting rather than as a complement to it. The distinction runs through AI Pre-Release Vetting: Ball's May 2026 position was that the federal government's ad-hoc pre-deployment review of Anthropic's Mythos model constituted a de facto licensing regime without legal basis, and that CAISI-led narrow-domain evaluations plus state-authorized IVOs were the defensible substitute (Claude Mythos Preview; Anthropic). On that framing the IVO question is not whether frontier models are reviewed but who holds the reviewing authority and under what legal instrument.

The mechanism is also distinct from Autonomy Certificates, a proposal for a third-party body to issue certificates communicating an agent's designed level of autonomy: certificates describe a design property to other developers and regulators, whereas IVO certification attests that a developer's risk-management practice meets a standard.

Verifier access

None of the four instruments specifies what a verifier may run, read or ask. The most concrete answer available comes from a different quarter: METR's July 28, 2026 proposal for third-party investigation of misalignment incidents, which names four categories of access an outside investigator would need — the ability to run all models involved, full transcripts or environments that reproduce the incident, employee interviews across security-and-infrastructure staff, training-data-and-RL staff, and anyone involved in an internal investigation, and the ability to run prompted classifiers over the training data — plus an inference budget, sufficient time, and AI tooling. METR tiers these: a transcript-and-interview-only investigation "would provide more limited assurance and be insufficient to answer many important questions about the root causes," while establishing whether a behavior traces to reinforcement-learning trajectories may require training-data ablations and access to intermediate checkpoints and previous related models (How independent researchers could investigate AI propensities after misalignment incidents (METR, July 2026)).

METR's proposal concerns post-incident investigation rather than periodic certification, and it does not address who selects or pays the investigator — the question the IVO debate turns on. Its sharing protocol nonetheless supplies a design element absent from the instruments: a redaction summary in which the investigator states how redaction limited the conclusions it could publicly substantiate, alongside full disclosure of the engagement terms — redaction terms, access provided, time and personnel, and agreed scope.

Open questions

  • Whether Virginia's IVO instrument is an enacted bill or a study directive, given the conflicting accounts recorded above.
  • Whether California SB 813 survived the 2025–26 session, and whether the tort-liability shield Weil describes was present in an earlier version of that bill or in a companion measure.
  • Whether the five-IVO Connecticut pilot appears in the enacted text of Connecticut SB 5 — Broad AI law (frontier reporting + ADMT + AI companions + sandbox) or in a separate Connecticut instrument; the SB 5 page does not currently record the provision.
  • What "adequacy audit" means operationally under the FRONTIER Act, and what access an IVO would receive.

Relationships