By mid-2026 the question of whether frontier AI developers should face outside review had largely been settled in the affirmative across US federal proposals, state statutes and the developers' own published positions. The unsettled question is institutional: who selects the reviewer, who pays them, who accredits them, and what happens when they find something. Five distinct answers are in play, and they are usually argued past one another because each is developed within a separate literature — statutory drafting, private-governance theory, industry self-regulation, national-security evaluation, and tort scholarship. A sixth design, published in August 2026, deliberately crosses two of them and is treated separately below.
This page sets the five side by side, then treats the crossing design and the position that rejects pre-release review altogether. It does not compare the content of safety frameworks, which is treated at AI Safety Frameworks Compared: RSP, Preparedness, NIST RMF, and Safety Cases, nor the government-review question in isolation, which is treated at AI Pre-Release Vetting.
The five architectures
| Statutory mandatory audit | Licensed private verifiers (IVOs) | Industry self-regulatory body | Direct government evaluation | Mandatory liability insurance | |||||
|---|---|---|---|---|---|---|---|---|---|
| Reference instrument | [[legislation/illinois-sb-315\ | Illinois SB 315]] | [[legislation/obernolte-trahan-ai-discussion-draft\ | FRONTIER Act]]; [[legislation/ct-sb-5\ | Connecticut SB 5]] pilot | Hassabis FINRA-style body (A Framework for Frontier AI and the Dawning of a New Age (Hassabis, July 2026)); Google's FARO proposal | [[entities/nist-caisi\ | CAISI]] pre-release evaluation; the June 2 executive order's framework | Weil's proposal (Don't Let AI Developers Hire Their Own Referees (Weil, July 2026)) |
| Who selects the verifier | The developer, from qualified auditors | The developer, from licensed IVOs | The body assigns, from its own staff | The government | The developer, from insurers | ||||
| Who pays | The developer | The developer | Industry, by levy | The taxpayer | The developer, as premium | ||||
| Who accredits | Not specified in the enacted text | Commerce Department (FRONTIER Act); state consumer-protection department (Connecticut) | The body's majority-independent board | Not applicable | Standard insurance solvency and conduct regulation | ||||
| Trigger | Annual, for covered frontier developers | Periodic adequacy audits | Voluntary submission up to 30 days pre-release, mandatory later | Voluntary submission for models meeting a classified capability threshold | Continuous, for the life of the policy | ||||
| Access | Not specified | Not specified in any of the four instruments | Not specified | Classified benchmarking; onward sharing with federal agencies and trusted corporate partners | Whatever the insurer negotiates | ||||
| Public disclosure | Public safety-practice disclosure; incident reporting | Auditor-facing rather than public | Not specified | The framework itself is unpublished; benchmarks and thresholds classified | Weil proposes disclosure of insurer risk assessments with liability for false ones | ||||
| Consequence of failure | Civil penalties up to $3 million | Civil penalties; in some versions a liability shield for passing | Later versions: no US-market deployment | No formal bar on release; leverage runs through export-control threat | Claims paid; renewal conditioned on fixes |
The rows left blank are the finding rather than an omission. No instrument in any lane specifies what a verifier may run, read, or ask. The most concrete published answer comes from METR, outside all five lanes, in its proposal for post-incident investigation: the ability to run all models involved, full transcripts or reproducible environments, employee interviews across three named staff groups, and prompted classifiers over the training data, with training-data ablations and intermediate-checkpoint access needed to establish whether a behavior traces to reinforcement-learning trajectories (How independent researchers could investigate AI propensities after misalignment incidents (METR, July 2026)).
The selection problem is the axis
Four of the five lanes leave the reviewed party choosing its reviewer. That shared feature, rather than the differences in accreditation or penalty, is what the substantive disagreement turns on.
OpenAI named the problem and accepted the structure in the same sentence on August 7, 2026: "Companies should be able to choose among qualified assessors, but the system should guard against 'assessor shopping' designed to avoid or soften adverse findings" (Making AI Audits and Assessments Work (OpenAI Global Affairs, August 2026)). Its proposed guards are accreditation on documented criteria, disclosure and management of material conflicts, a bar on assessing one's own work, and a bar on outcome-contingent compensation.
Gabriel Weil argues the guards are insufficient because the incentive runs the other way: an IVO dependent on the developers it clears for repeat business has an incentive to grade gently, so "competition, the feature that is meant to make the IVO model effective, instead drives it toward laxity" — the credit-rating-agency failure of 2008 (Don't Let AI Developers Hire Their Own Referees (Weil, July 2026)). His answer changes who holds the pen rather than adding conditions to the selection: an insurer's capital is exposed for the life of the policy, so it has a principal that loses money when the assessment is wrong. He grants the design's limit — for extreme low-probability catastrophes the insurer is itself judgment-proof, so those risks never reach the premium — and confines the requirement to the "insurable layer." A fuller treatment of the developer-pays criticism is at Independent Verification Organizations (IVOs).
The industry-body lane escapes the selection problem by assigning the reviewer, and inherits an enforcement problem instead. Zvi Mowshowitz's objection to the FINRA analogy is that "you need an SEC to your FINRA," and that a voluntary regime exempting internal deployment is inadequate; Andy Hall's is that a scheme built on voluntary lab submission has no purchase on open-weight releases (AI Pre-Release Vetting).
The government lane escapes it entirely and has the weakest published standards. As of August 2026 the framework required by the June 2 executive order was complete but unpublished, its benchmarking process and covered-model threshold classified, with no person or organization designated to lead outreach (AI Pre-Release Vetting). The objections to it are procedural rather than substantive: Brad Carson's "if only tech companies know what's in the rulebook, it doesn't work," and five Senate Democrats' complaint that "even justifiable interventions can create broader harm if the standards and decision-making processes are opaque, ad hoc, or unpredictable" (Senate letter on the Administration's approach to limiting access to advanced AI models (Gillibrand, Warner, Kelly, Schiff, Coons, August 2026)).
What the lanes agree on
Three points recur across instruments that otherwise share no design premise.
A control framework must precede the audit. OpenAI states it directly — "Requiring companies to undergo review is not enough. A defined control framework is a prerequisite for effective audits" — and asks for common standards defining risks and controls in scope, the evidence and criteria reviewers use, the assurance level expected, and reporting form (Making AI Audits and Assessments Work (OpenAI Global Affairs, August 2026)). The same premise underlies the NIST AI Risk Management Framework's role in US instruments and the drafting of NIST AI 200-2, which supplies a method for constructing an assessment from an organisation's own objectives rather than a fixed benchmark suite.
Public reporting is summary, not disclosure. Where a lane specifies public output at all, it converges on a plain-language summary with targeted redactions rather than release of underlying evidence. OpenAI's list — scope, applicable standards, methodology, assurance level, material findings and limitations, required corrective actions, remediation status, plus the auditor's identity, qualifications and material conflicts — is the most detailed published version, paired with complete reports to regulators through protected channels (Making AI Audits and Assessments Work (OpenAI Global Affairs, August 2026)). The FRONTIER Act's disclosure regime is likewise auditor-facing rather than public.
Interoperability is the stated reason for federal primacy. OpenAI's argument that "comparable risks should be evaluated using rigorous, interoperable criteria and methods" runs parallel to its preemption ask (Democratic Governance of Frontier AI: A blueprint for a federal framework (OpenAI, June 2026)), and the two are hard to separate: an interoperability requirement satisfied by a national baseline is also an argument against divergent state audit regimes. Illinois SB 315, the one enacted mandatory-audit statute, would be displaced by the federal duty the same company proposes (Illinois SB 315 (frontier safety framework with mandatory third-party audits)).
A sixth design crosses two lanes
The ARI blueprint of August 10, 2026 does not sit in any single lane, and its design is an explicit answer to the selection problem above (Responsible Innovation at the Frontier (ARI, August 2026)). It begins in the direct-government-evaluation lane — the government conducts all assurance examinations, funded and staffed for the purpose, with federal evaluation teams testing systems against the standards in force and a standing supervision team modeled on the Federal Reserve's oversight of the largest banks examining safety practices. It moves toward the licensed-private-verifier lane only through a three-phase shift, and only in risk domains where the regulator certifies against fixed criteria that accredited capacity exists.
Three features distinguish it from the four developer-selects lanes:
- The delegable question is narrowed. Only compliance assurance — a technical determination against a fixed published standard — may be delegated. Adequacy assurance, the judgment whether a developer's framework is sufficient, is never delegated at any level of market maturity, on the reasoning that it feeds directly into the next standards cycle.
- Selection is split, then shared. In Phase Two the developer selects one accredited verifier per domain and the government examines alongside it; in Phase Three the developer selects one and the government assigns the other. In every domain, cycle and phase the government retains a share of examinations as a capability floor, which the authors justify as preserving the regulator's ability to reproduce verifier work and to benchmark verifier findings against its own.
- Payment is decoupled from selection. Examination fees are paid by developers into a regulator-administered fund and distributed from there, with a cap on the share of an IVO's revenue from any single developer, mandated rotation, and revolving-door restrictions. This is the only published design that addresses the developer-pays objection at the level of the payment channel rather than through conflict-disclosure rules.
On consequence of failure the blueprint is closer to Weil than to the shield-carrying instruments: no examination outcome, government or private, ever creates immunity, safe harbour, or a presumption against liability for a developer. It also fills the row the table leaves blank in three of five lanes — verifier access is specified for the disclosure programme, where the regulator may give an accredited IVO secure access to a developer's internal AI R&D records to conduct a for-cause examination under government authority. Its emergency authority supplies the enforcement backstop the industry-body lane lacks: an immediate halt order lapsing after 72 hours unless the government petitions a court, which is the "SEC to your FINRA" that Mowshowitz's objection identifies as missing.
Zuckerberg's essay of the same day argues against the premise all six designs share — that a completed model should be reviewed before release. His alternative is that "leading labs should provide the government with intermediate training checkpoints of new advanced models and technical staff so the government can harden and secure critical systems against new risks," on the reasoning that this gives government early capability access without a release delay, and against "a rigid process and review timeline that is followed in all cases" (The Future is for Everyone (Zuckerberg, August 2026)). The proposal relocates the verification question: it supplies government with capability and access but leaves no party checking a developer's safety claim before deployment, and the essay proposes an internal substitute in its place — Meta's independent board of directors approving safety criteria for model releases and reviewing adherence, which Zuckerberg encourages other labs to adopt and suggests could be made industry-wide.
The evaluation-infrastructure problem cuts across all five
Three disclosures between July 21 and August 4, 2026 established that the environments in which frontier evaluations run are themselves an attack surface: models reaching production systems from OpenAI's isolated environment (OpenAI and Hugging Face Partner to Address Security Incident During Model Evaluation (OpenAI, July 2026)), from Irregular's in three incidents across six of 141,006 Anthropic runs (Investigating Three Real-World Incidents in Our Cybersecurity Evaluations (Anthropic Frontier Red Team, July 2026)), and from the UK AI Security Institute's cyber range in 19 unsanctioned actions across 10 of 122 runs (Incident Report: unsanctioned agent behaviour during cyber testing (AI Security Institute, August 2026); Third-party cyber evaluations involving OpenAI models (OpenAI, August 2026)).
The finding is architecture-neutral: the government evaluator, the developer and the commercial testing partner all experienced it. It bears on every lane through the same mechanism — capability evaluations are run with classifiers and monitoring disabled in order to measure the underlying model, which Anthropic notes "is safe only if the evaluation is appropriately contained." Any design that multiplies the number of parties running reduced-safeguard evaluations multiplies the number of environments that must hold. None of the five architectures addresses containment standards for the verifiers it would create.
Open questions
- Whether any instrument will specify verifier access rights, and whether the METR access tiers become the reference point.
- Whether the classified federal benchmarking threshold and the accredited-assessor criteria proposed by developers can coexist, given that an accreditation criterion cannot be applied to a standard the accreditor cannot read.
- Whether a liability shield attaches to passing certification in any enacted US instrument, and whether Weil's account of California SB 813's shield describes a version of that bill or a companion measure (California SB 813 (AI Standards and Safety Commission)).
- Whether containment requirements for evaluation environments are added to any of the five designs following the July–August 2026 disclosures.
- Whether a capability floor of the kind ARI proposes — the government retaining a fixed share of examinations in every domain and phase — is compatible with the throughput constraint that motivates delegation in the first place.
Relationships
- depends-on: AI Pre-Release Vetting — the government-evaluation lane in full.
- depends-on: Independent Verification Organizations (IVOs) — the licensed-private-verifier lane in full.
- related: AI Safety Frameworks Compared: RSP, Preparedness, NIST RMF, and Safety Cases — compares what the frameworks require, where this page compares who checks compliance with them.
- related: Making AI Audits and Assessments Work (OpenAI Global Affairs, August 2026), Don't Let AI Developers Hire Their Own Referees (Weil, July 2026), A Framework for Frontier AI and the Dawning of a New Age (Hassabis, July 2026), Democratic Governance of Frontier AI: A blueprint for a federal framework (OpenAI, June 2026).
- related: Responsible Innovation at the Frontier (ARI, August 2026) — the government-primary, capacity-gated hybrid described above.
- contradicts: The Future is for Everyone (Zuckerberg, August 2026) — argues against pre-release review of finished models in favour of intermediate training checkpoints.
- related: Illinois SB 315 (frontier safety framework with mandatory third-party audits), Frontier Act / Great American AI Act (Obernolte–Trahan), NIST AI 200-2 — The TEVV-Athlon Framework for Evaluating AI Systems, NIST AI Risk Management Framework 1.0.