AI Policy Wiki
Dashboard

Anthropic's Advanced AI Framework (June 2026)

high confidence · updated 2026-07-25

Anthropic's proposed near-term government framework for frontier AI catastrophic risk, in two parts: obligations on covered frontier developers (safety framework, six-monthly risk reports, system cards, 15-day critical-safety-incident reporting, mandatory independent evaluation, a security program, and enforcement with civil penalties scaling to global annual revenue), and cross-government investments in biological and cyber societal resilience. Covers developers training above 10^25 FLOP with over $500M AI revenue or over $1B AI R&D spend, across four enumerated risks.

Anthropic's Advanced AI Framework, published June 2026, is the company's proposal for "what we think governments should do in the near term about the most serious risks from frontier AI." It is the legislative-proposal companion to Dario Amodei's essay "Policy on the AI Exponential", and is written primarily for the US federal government, with the note that other jurisdictions should tailor the recommendations to their own capacity and authority.

The document has two parts: obligations on frontier developers, and cross-government investments in societal resilience "so that a biological or cyber attack is harder to carry out and easier to recover from, wherever the capability originates." Anthropic states its own confidence unevenly — "We are more confident about some parts of this framework than others, and recognize it draws from both existing and novel concepts" — and defends the exercise on the ground that "the cost of waiting for perfect policy is having none during a critical period."

Part 1 — Frontier developer obligations

Scope

ElementProvision
Covered DeveloperMust meet both: develops models requiring more than 10²⁵ training FLOP; and earns more than $500 million in annual AI-derived revenue (inflation-adjusted) or spends more than $1 billion per year on AI R&D.
Enumerated Risks(a) biological weapons; (b) offensive cyber operations; (c) loss of control; (d) automated research and development in key domains that could accelerate or amplify (a)–(c).
Catastrophic Risk"A foreseeable and material risk that a Covered Developer's development, storage, use, or deployment of a covered model will materially contribute to significant death, injury, or damage" — expressly aligned with the definition in [[legislation/california-sb-53California SB 53]].
Critical Safety Incident(a) unauthorized exfiltration, or unauthorized deliberate malicious modification, of a covered model's weights; (b) harm from a materialized Catastrophic Risk; (c) loss of control causing death or bodily injury; (d) a model using deceptive techniques against the developer to subvert controls or monitoring outside an evaluation designed to elicit that behavior, in a manner demonstrating materially increased Catastrophic Risk.
Periodic reviewAn Agency should review Covered Developer criteria at least annually with input from industry, academia, government, and civil society, paired with "procedural, judicial, and structural" safeguards against overbroad discretion.

Anthropic notes that the FLOP needed to train dangerous models may fall over time, and that "it may make sense to introduce a threshold based on capabilities rather than simply on training costs." See Risk-Based AI Regulation.

Preemption

The framework takes a specific and restrictive position: Congress "should not preempt state law unless it enacts a rigorous federal regime that meets or exceeds the strongest measures proposed in this framework." Even then, preemption should be limited to "the specific frontier governance functions Congress has expressly chosen to occupy, such as catastrophic-risk testing of covered models, evaluator licensing, or closely related reporting obligations."

Four further limits follow: federal law should not be read to occupy the field of AI regulation, displace state law by implication, or preempt state statutory or common-law claims; any preemption should be construed narrowly with ambiguity resolved in favor of state authority; compliance with a federal regime "should not itself confer immunity, a safe harbor, or a presumption against liability" under state law; and absent a strong federal law, "states should retain full authority to legislate."

Transparency

Anthropic states that "transparency alone is not sufficient for advanced AI," positioning it as a record rather than a remedy. Four instruments are proposed:

Safety framework. Covered Developers must develop, publish, and follow a framework identifying the covered models it applies to; describing how each Enumerated Risk is evaluated, including standards relied on, capability evaluations performed, and mitigations applied; identifying the corporate officer primarily accountable for implementation; and describing the process for modifying the framework. Annual certification of compliance to the Agency is required, with civil penalties for material misrepresentation. See Responsible Scaling Policy (RSP).

Risk reports, at least every six months, giving an overall assessment of each Enumerated Risk. Minimum contents: a representative summary of relevant capabilities across deployed models with material changes since the prior report, covering all externally served models "as well as any substantially riskier models the developer deploys internally"; the threat models tracked, how observed capabilities map to them, and the mitigations addressing each; and residual risk in each category after mitigations, including internal deployments. The framework sets an explicit sufficiency test for public disclosure: "It should be possible to understand most of the reasoning behind the risk assessment from public information, and to reach a similar overall conclusion about the level of risk as one would reach with access to all the information the developer has." It notes a higher frequency than six months may be needed if AI progress accelerates.

System cards, published when deploying a covered model for general access that, in any Enumerated Risk category, is either materially more capable than any previously deployed covered model, or is deployed under materially weaker safeguards than a previously deployed model of comparable or greater capability. Contents: testing methodologies and results; capabilities, limitations, and intended and observed behaviors; and how internal or external deployment changes the risk assessment relative to prior system cards and risk reports.

Critical Safety Incident reporting to the Agency within 15 days of discovery, or of facts giving reasonable belief that such an incident occurred. Reports should be shared with relevant federal agencies and national laboratories, and exempted from public-records disclosure laws consistent with existing state law.

Redactions. Developers may redact trade secrets, material that would compromise public safety, model security, or national security, or material withheld to comply with law — "only as necessary," with the nature of each redaction noted in the published version.

Independent evaluation

The framework's stated premise is that "self-assessment is not enough," coupled with the acknowledgment that "a mature independent evaluation ecosystem does not yet exist." Within six months of enactment, Covered Developers would regularly engage at least one qualified independent evaluator.

Evaluator access rights: an unredacted version of the most recent risk report and system cards, access to the developer's most capable models, and the opportunity to ask and receive reasonable responses about models, likelihood of catastrophic risks, and safeguards. Independence is defined as having no financial interest in the developer and no major conflicts of interest with the individuals conducting the review.

Evaluators would publish a review of the developer's most recent risk report addressing four things: the adequacy of its information, its analytical rigor, "the appropriateness and materiality of any redactions made to the public version," and whether the evaluator disagrees with any key claims, "especially its overall assessment of the level of risk for each Enumerated Risk category." Evaluators would be bound not to copy, retain, or disclose confidential information — naming confidential intellectual property, national-security matters, and proprietary details such as model architecture, cost, or size — but "beyond that should generally not be restricted in what they can publish, including concerns about the risk report or the developer's conduct in connection with the review process."

Four ecosystem measures accompany this: publishing standards for evaluators and exploring a licensing system; government or pooled funding so evaluators remain financially independent of any developer; independence certification covering funding and remuneration, conflict-of-interest and equity-ownership disclosures, board participation, and employment, with possible post-employment cooling-off periods; and a named remedy for "evaluator shopping" — the risk that companies seek whichever evaluator asks least of them — under which agencies could rate evaluators on predefined criteria such as rigor of public reasoning, with highly rated evaluators randomly assigned to developers in high-stakes cases. The framework adds that a government evaluation function "could eventually supplement or even substitute for an independent evaluator ecosystem."

Security

The stated principle is that "a model that is safe to deploy is not safe if it can be stolen or quietly copied." Covered Developers would maintain a security program covering the full development environment — model weights, training and inference infrastructure, trusted-partner access controls, and internal development processes — robust to external and insider threats and scaled to the consequences of compromise; describe it at a general level publicly and in detail to the Agency on request; report known model-extraction and distillation attacks along with detection and prevention measures to the Agency and to other Covered Developers; and conduct regular red teaming and penetration testing covering weights, algorithmic secrets, training infrastructure, and insider threats, informing the Agency of finding categories and remediation status. Independent evaluators and any third party granted privileged or pre-deployment access — expressly including access to versions with safety mitigations reduced or removed — would maintain commensurate security protections.

Enforcement

Three provisions are proposed to hold "under any enforcement model": a prohibition on intentionally false or materially misleading statements about safety-framework compliance, required evaluations, or required disclosures; civil penalties for failure to conduct required evaluations, publish a safety framework, or report incidents, escalating with repeated violations and scaling with global annual revenue; and whistleblower protections requiring anonymous internal channels for employees and contractors, prohibiting retaliation against good-faith reporters and contractual restrictions on reporting.

Beyond that, the framework says "there should also be a way to block or deter deployment of models that pose significant catastrophic risks," concedes "much room for debate," and presents options rather than a single design — suggesting policymakers "could begin with a lighter-touch model and revisit that choice as model capabilities advance and the independent evaluation ecosystem matures."

Agency review would identify four violations: required system cards, risk reports, or independent evaluations not completed and published; evaluation not done by a sufficiently disinterested or qualified evaluator; evaluator lacking sufficient time or access for a high-integrity evaluation; or the risk assessment or independent evaluation finding that deployed covered models pose significant risk of catastrophic harm in an Enumerated Risk category even accounting for active safeguards.

Remedies: fines for deploying models with inadequately mitigated risks; prohibitions on deploying further covered models until violations are corrected; and "in extreme cases, requirements to restrict usage of, and access to, already-deployed models."

Safeguards against overreach, offered as mixable options: court enforcement, under which the Agency may pursue remedies only by suit — with provisional remedies permitted for imminent catastrophic risks, reversed if a court does not uphold them within a set period, or fast-tracked litigation; cabined discretion, under which the Agency may act only on the enumerated violations and not on its own assessment of risk; and consistent treatment with judicial review, holding all covered models of equivalent capability to the same standards, barring advantage or disadvantage on grounds unrelated to the evaluation record or Enumerated-Risk capabilities, and providing expedited judicial review of remedies.

Part 2 — Societal resilience measures

The second part addresses "how society can prepare to withstand the threats — particularly biological and cyber — that advancing AI capabilities may accelerate or enable." Anthropic's stated rationale for treating this as urgent is timing: surveillance systems, stockpiles, hardened infrastructure, and response capacity "take years to build and cannot be stood up in a crisis." It notes these are established fields with existing institutions and live debates predating frontier AI, offers the recommendations "as priorities and directions rather than finished policy designs," and observes that many are worth making regardless of how AI develops.

Biological resilience

The framing distinguishes biological threats by their dynamics: a released agent "replicates, spreads person-to-person, and evolves under selection pressure," so a small breach can become large without further attacker action and the agent can adapt to evade countermeasures. Measures are organized as layers.

Prevention. Modernize and unify national biosafety and biosecurity rules, "many of which predate synthetic biology, directed evolution, and AI-assisted design tools." Extend enforceable standards to privately funded research, strengthen institutional review committees, and require personnel vetting for access to the most dangerous pathogens. Require gene synthesis providers and benchtop synthesizer manufacturers to screen requested sequences for known or predicted hazards and verify customers seeking sequences of concern. Dedicate intelligence and law-enforcement resources to disrupting bioweapons pursuit and attributing state-level activity, and establish two-way CBRN threat-intelligence channels between governments and AI developers, supported by legal safe harbors, information-sharing protections, and antitrust carve-outs, with independent stress testing of biosecurity safeguards.

Detection. Establish, fund, and maintain pathogen-agnostic biosurveillance at national, subnational, and international levels for actionable early warning. Build and exercise federal microbial-forensics and attribution capability to "underwrite deterrence-by-denial," maintaining reference databases, accredited laboratories, standing rapid-deployment forensic teams, and international channels for cross-border evidence sharing.

Preparedness and response. Harden public-health and emergency-response infrastructure, workforce, and logistics, testing plans against worst-case biotechnology-enabled outcomes in live drills. Stockpile pandemic-grade respiratory protection — "especially reusable respirators" — for the essential workforce with surge manufacturing capacity. Support standards, R&D, and investment in airborne-transmission suppression in critical infrastructure (utilities, schools, hospitals, transit hubs, government facilities). Use AI-accelerated drug discovery and protein design to shorten countermeasure timelines, funding partnerships between AI developers and biodefense institutions. Increase investment in broad-spectrum antivirals, which reduce the pressure to rapidly characterize a pathogen. Invest in adaptable platforms and manufacturing for vaccines, therapeutics, and diagnostics against novel or engineered pathogens, stockpiling precursors and maintaining pre-established clinical-trial protocols. Require binding after-action reviews after any nationally significant biological incident, with public reporting, statutory deadlines for corrective action, and independent oversight of unimplemented findings.

See AI Biosecurity, CBRN Uplift.

Cyber resilience

The framing is economic: frontier AI "is shifting the economics of cyber offense, accelerating vulnerability discovery and exploit development at a scale that will spread well beyond the handful of developers covered by Part 1," and the aim is to ensure "the same technology that lowers the cost of attack also lowers the cost of defense."

Prevention. Fund sustained maintenance, security auditing, AI-assisted vulnerability remediation, and migration to memory-safe languages for the open-source and legacy software underpinning critical infrastructure and the internet. Accelerate adoption of phishing-resistant authentication and content-provenance standards. Fund forward-deployed engineers, shared regional security operations centers, and managed security services for under-resourced critical-infrastructure operators — naming water and wastewater utilities, municipal governments, school districts, regional hospitals, and rural electric cooperatives — prioritizing "the basics with the highest payoff against known-but-unpatched vulnerabilities": asset inventories, prompt patching, and reducing internet exposure of operational technology via data diodes.

Measurement and situational awareness. Direct the national AI safety institute or equivalent to develop and maintain cyber capability evaluations in partnership with the intelligence community and national-laboratory operational-technology testbeds, issuing capability guidance as thresholds are crossed. Establish a dedicated threat-intelligence function fusing developer abuse monitoring, government reporting, and incident-response telemetry into a single picture of who is misusing or distilling frontier cyber capabilities, with short-term data retention encouraged. Provide legal safe harbor for AI developers to share defensive research, vetting practices, jailbreak data, distillation defenses, and threat intelligence with each other, the government, and the wider ecosystem.

Defender advantage. Distribute best practices on hardware-root-of-trust network isolation to prevent lateral movement, and on using ML and AI to detect and remediate breaches "within minutes or hours instead of days or weeks." Assist developers in co-developing safeguards — vetting, user verification, misuse and distillation safeguards, threat-intelligence sharing — enabling responsible release of cyber capabilities to large numbers of defending organizations. Use industrial-base and defense-production authorities to build a strategic reserve of long-lead-time operational-technology hardware (protective relays, remote terminal units, programmable logic controllers) for grid and water operators. Fund the national cybersecurity agency to map the most critical AI and software supply-chain nodes, and to handle 10–100 times the current disclosure volume with sub-seven-day turnaround, with the relevant standards body revisiting disclosure timeline norms for the AI era. Establish binding patch-deployment frequency for critical-infrastructure operators calibrated to sector-specific constraints, funding AI-assisted patch generation, backporting, and validation "so that faster timelines are achievable rather than aspirational."

Security modernization. Require vendors of software and connected devices used in critical infrastructure to publish support lifecycles and end-of-life dates; fund operators to inventory and isolate systems that can no longer receive updates; and establish a funded replacement program modeled on prior national-security-driven equipment replacement, prioritized by exposure and consequence. Fund research transcending find-and-patch — formal verification, runtime patching, polymorphic defense — with a requirement that it "demonstrate value beyond what frontier models can already do, rather than reproducing their capabilities."

See AI and Cybersecurity, Autonomous cyber-agents.

Loss of control and automated R&D

The framework is explicit that this half of its own risk taxonomy has no developed resilience agenda: the societal resilience agenda here "is less mature than it is for biological and cyber risks, and we believe it needs much more active work across the field." The directions it names are the capacity to detect and respond to AI systems acting outside their developers' control, and infrastructure for containing or shutting down such systems, with a commitment to share more as Anthropic's understanding matures. See AI Autonomy Risk.

Relationships