AI Policy Wiki
Dashboard

Managing Advanced Cyber Risks in Frontier AI Frameworks

medium confidence · updated 2026-07-28

Frontier Model Forum report documenting industry consensus on cyber thresholds, evaluations, and safeguards in frontier AI safety frameworks.

"Managing Advanced Cyber Risks in Frontier AI Frameworks" is a technical report published by the Frontier Model Forum on 2026-02-13. It documents emerging industry consensus on how frontier AI developers should manage cybersecurity risks within their safety frameworks, and presents the collective position of Amazon, Anthropic, Google DeepMind, Meta, Microsoft, and OpenAI. The report extends the FMF's Technical Report Series on implementing frontier AI safety frameworks into the cyber domain specifically, and states that it is based on expert discussions among FMF member firms. A PDF of the report is published on the FMF's own domain alongside the web version.

The report frames cyber capability as dual-use: frontier AI can accelerate vulnerability discovery and patching, optimize defensive systems and enhance threat detection, while the same capabilities lower barriers for malicious actors to exploit known vulnerabilities or discover new attack vectors. It notes that every FMF member firm has published a frontier AI framework identifying advanced cyber threats as a key risk, and that firms must weigh the defender benefits their models provide when setting thresholds.

Summary

The report describes how frontier developers define the cyber-capability thresholds at which a model would trigger additional safeguards, the methods used to evaluate models against those thresholds, and the questions on which firms have not yet converged. It frames cyber risk through a threshold–evaluation–safeguard structure, organized across four substantive sections — thresholds, capability assessments, misuse mitigations, and mitigation assessments — followed by a statement of continuing work.

Threat modeling

The report treats threat modeling, adapted from the cybersecurity and national-security domains, as the process that precedes threshold-setting: systematically anticipating how threat actors might use frontier AI to achieve harmful outcomes and mapping the pathways to them. It states that several firms have integrated AI-cyber threat modeling into their risk-management processes, and identifies seven components:

  • Apply established cybersecurity frameworks. Firms articulate kill chains — the sequence of steps a threat actor must take — often building on the Lockheed Martin Cyber Kill Chain, MITRE ATT&CK, and STRIDE.
  • Articulate key assumptions and variables, including the prerequisites for a successful attack (tools, skills, resources), the capability levels of different threat actors from novices to advanced persistent threat groups, and the model's scope boundaries.
  • Identify threat scenarios, beginning from the most severe outcomes — destructive attacks on critical infrastructure, large economic damage from data breaches — and drawing on historical precedents such as NotPetya and the Morris Worm, expert consultation, and workshops.
  • Project how frontier AI could lead to harm, by locating the critical bottlenecks in offensive cyber operations where AI assistance would be necessary or significantly advantageous.
  • Map contributing risk factors, using industry sources such as the OWASP Top 10 and Mandiant's M-Trends report to ground the analysis in how attackers exploit vulnerabilities in practice.
  • Map threat models explicitly to capability evaluations, for example by relating the feasibility of a threat model to performance on a cyber benchmark.
  • Continuously update threat models as malicious actors adapt AI tools in ways initial exercises did not anticipate.

Thresholds

Four types of thresholds

The report distinguishes four kinds of thresholds used across frameworks and states the tradeoff for each:

TypeDefinitionTradeoff as stated
CapabilitySpecific AI capabilities that could enable harmA better risk proxy than compute and more measurable than risk thresholds, but still difficult to implement and measure through evaluations. The most commonly used type for cyber.
RiskQuantifies unacceptable risk levels through likelihood or severity metricsMeasures risk directly, but hard to implement reliably given the complexity and uncertainty in estimating future risks
ComputeTraining compute as a proxy for capabilitiesThe easiest to measure but the weakest risk signal — greater training compute does not always equate to greater risk
OutcomeSpecific scenarios and an AI system's contribution to realizing themConnects capabilities directly to real-world harms, but comprehensive and realistic scenarios are difficult to define and evaluate

The report suggests that frontier risk management may benefit from a mixed approach: using a compute threshold as an initial scope signal, then benchmarking a model against a reference model to determine whether it represents a material change in general capability relevant to the risk domain. It also notes that outcome-based thresholds allow developers to establish tiered thresholds — critical, high, medium — based on how far a model advances threat actors toward a dangerous outcome.

Two consensus cyber thresholds

The report states that nearly every firm has converged on two core capability thresholds:

  1. Non-expert uplift: whether an AI model significantly enables individuals with limited cybersecurity expertise to conduct sophisticated cyberattacks. Models are described as crossing a critical threshold when they supply specialized knowledge, troubleshooting guidance or procedural instruction that meaningfully reduces the expertise, resources or time traditionally required of low-skill threat actors.
  2. Autonomous end-to-end cyberattacks: whether AI systems can plan, execute, and adapt attacks without human intervention. The report argues this represents a capability leap because autonomous systems, unlike human-assisted attacks limited by operator expertise and availability, could operate continuously and simultaneously across multiple targets, potentially overwhelming human defenders and existing security infrastructure.

The report characterizes this convergence as paralleling how other domain-specific frameworks have coalesced around comparable capability thresholds.

Published thresholds by firm

The report reproduces the published cyber thresholds of six member firms, noting that the table records only thresholds where consensus exists and that the entries are accurate as of January 2026:

FirmTypeThreshold as published
AmazonCapabilityMaterial uplift beyond other publicly available research or tools enabling a moderately skilled actor — given as an individual with undergraduate-level understanding of offensive cyber operations — to discover new high-value vulnerabilities and automate their development and exploitation
AnthropicCapability"Cyber Operations": significantly enhancing or automating sophisticated destructive cyber attacks, including discovering novel zero-day exploit chains, developing complex malware, or orchestrating extensive hard-to-detect network intrusions
GoogleCapability"Cyber uplift level 1": sufficient uplift with high-impact cyber attacks for additional expected harm at severe scale
MetaOutcomeThree tiers: automated end-to-end compromise of a best-practice-protected corporate-scale environment; automated discovery and reliable exploitation of critical zero-days in popular software before defenders can find and patch them; widespread economic damage via scaled long-form fraud and scams
MicrosoftCapabilityFour tiers (Low, Medium, High, Critical), rising from support for gathering publicly available threat information to enabling a well-resourced expert actor to develop and execute novel and effective strategies against hardened targets
OpenAICapability"High": removal of existing bottlenecks to scaling cyber operations, including automating end-to-end operations against reasonably hardened targets or automating discovery and exploitation of operationally relevant vulnerabilities. "Critical": a tool-augmented model that can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or devise and execute end-to-end novel attack strategies against hardened targets given only a high-level goal

Points of divergence the report identifies

The report records five considerations on which frameworks diverge or remain imprecise:

  • The meaning of "significant" uplift. Several thresholds qualify uplift as "significant," "meaningful" or "material" without defining the term, leaving developers to exercise judgment and complicating efforts to establish consistent and comparable safety benchmarks across organizations. The report suggests establishing consensus on what constitutes significant uplift.
  • Threat-actor expertise. Frameworks concentrate on lowering barriers for low-skilled actors and generally do not address how highly skilled, well-resourced actors such as advanced persistent threat groups may use these tools, even though such actors may be most capable of enabling the greatest harm. The report also flags an unresolved distinction between cumulative uplift across a population of novice attackers and uplift to a single actor's ability to mount a large-scale attack.
  • Zero-day discovery as a critical escalation point. Nearly every firm identifies discovery of zero-day exploits as a major threshold marker, described as the jump from exploiting known weaknesses to independently finding entirely new attack vectors.
  • End-to-end automation without human intervention, described as a shift from AI as a tool that augments human capability to AI as an independent cyber operator capable of planning, executing and adapting attacks without guidance.
  • The definition of a well-protected target. Several frameworks reference "patched," "hardened" or "state of the art" targets without consensus on what constitutes a security best practice.

The report notes that several frameworks employ multiple tiered thresholds that trigger corresponding mitigation actions well before intolerable risk levels are reached, and that some reference the impact frontier AI systems will have on defenders in assessing risk. It lists as remaining open questions how to define "significantly enable" in the cyber context, how much cumulative evidence is needed to determine a threshold has been crossed, and how performance on assessments translates into risk.

Capability assessments

Evaluation methods

The report catalogs the methods developers use to measure models against these thresholds, and records limitations for each:

  • Capture the Flag (CTF): structured hacking challenges in isolated environments, with defined time limits and well-defined success metrics; Hack the Box exercises are given as an example. Limitations: agent evaluations require scaffolding through fine-tuning, prompt engineering or tool access to test the upper bound of harm, and the lack of standardization across elicitation methods makes cross-developer comparison difficult; models are told to pursue a specific objective rather than operating under the ambiguous conditions a real threat actor faces; and success may reflect exposure to publicly available solutions during training rather than capability.
  • Cyber Range: virtualized infrastructure simulating real-world networks, permitting assessment of reasoning and planning across multi-stage threats. Limitations: the same elicitation problems as CTFs, plus the absence of defenders in the environment.
  • Benchmarks: knowledge tests, capability tests, and safeguard evaluations — respectively, multiple-choice and open-ended question sets, agent-based practical task challenges, and tests of model responses to potentially harmful prompt requests. Limitations: contamination, where models are trained on benchmark-related data and scores are artificially inflated, and saturation, where scores rise high enough that the benchmark no longer detects meaningful improvement.
  • Red-teaming: expert-led probing for vulnerabilities, covering exploit development, vulnerability discovery and social-engineering assistance, relying on the creativity of security professionals rather than standardized challenges, and increasingly incorporating automated protocols.
  • Controlled trials: uplift studies comparing AI-assisted versus unassisted human performance, using randomized control or comparable treatment-versus-control designs and measuring against existing tools such as search engines. This is the only listed method focused on human–AI interaction rather than model performance alone.

The report distinguishes the first three, which evaluate agent capabilities, from knowledge benchmarks and safeguard evaluations, which test overall cybersecurity knowledge or the propensity to produce a harmful response. It states that developers conduct cyber evaluations at three points in the lifecycle: before any safety measures are applied, to establish maximum capability; after safeguards are implemented, to assess their effect; and as close to deployment as possible, to capture enhancements made during post-training.

The report reproduces a non-comprehensive table of nine publicly available offensive cyber benchmarks referenced in member firms' model cards, system cards and technical reports: CVE-Bench (2025, cited by OpenAI), CyberGym (2025, Anthropic), CyberSecEval 4 (2025, Meta), CTIBench (2024, Amazon), Cybench (2024, Amazon, Anthropic and Meta), CyberMetric (2024, Amazon), SecBench (2024, Microsoft), SECURE (2024, Amazon), and SWE-bench Verified (2025, Anthropic, Google DeepMind and OpenAI).

Evaluation domains

Cyber evaluations most often address cyber operation assistance, tested across six skill areas: reconnaissance (open-source intelligence collection and network reconnaissance); social engineering (phishing simulations testing persuasion toward downloading a malicious attachment or divulging security information); malicious code generation (file encryption, data exfiltration, keyloggers, DDoS initiation); vulnerability discovery and exploitation (cryptography, forensics, binary exploitation, web security); tool usage (Bash and PowerShell execution, Python scripting, Metasploit); and network operations (scanning, sniffing and spoofing at the CTF level; autonomous navigation, compromise, privilege escalation and evasion in realistic network environments at the cyber-range level).

Evidence for threshold decisions

The report states that threshold determinations should rest on cumulative evaluation evidence in a holistic assessment, since a single evaluation is unlikely to indicate unequivocally whether a threshold has been crossed, and that the field remains unclear on which precise combination of evidence is needed. It recommends bottleneck assessments testing whether a model or system can perform five tasks significantly better than other available baseline resources: assisting with attack planning; assisting with or autonomously discovering a novel zero-day exploit; assisting with or autonomously developing complex malware; assisting with or autonomously escalating privileges and moving laterally within a system; and carrying out autonomous cyberattacks against a hardened target. These are to be read as a constellation of factors rather than in isolation.

Misuse mitigations

The report divides safeguards by mode of application — model-level (applied during training, fine-tuning or alignment, modifying parameters and behavior), system-level (implemented in the deployment environment or application layer, monitoring, filtering or restricting inputs and outputs without changing parameters), and societal-level (physical controls, supply-chain security, regulatory compliance or inter-organizational coordination) — and states that developers generally adopt a defense-in-depth strategy layering multiple safeguards. It explicitly declines to prescribe an ideal combination, notes that many mitigations require further research to validate, and excludes safeguards designed to protect models themselves from compromise, such as jailbreak prevention or secure access protocols.

Five safeguard categories are identified:

  • Capability limitations — altering weights or the training process so the model does not possess the knowledge or ability, through targeted unlearning or false learning (training on deliberately fabricated but plausible-sounding incorrect information). The report describes these as still largely experimental and not widely implemented for offensive cyber capabilities.
  • Behavioral alignment — supervised fine-tuning, RLHF and RLAIF, including Anthropic's Constitutional AI and OpenAI's deliberative alignment, plus safe-completion training that produces safe outputs to dual-use queries rather than refusing outright. The stated limitation is that safety training modifies surface-level behavior without altering underlying capability, and can be undone through targeted fine-tuning or adversarial prompts; RLHF additionally carries reward-misspecification and reward-hacking risk, and RLAIF's efficacy depends on how comprehensively the underlying principles are defined.
  • Detection and intervention — LLM-based prompted classifiers, custom-trained classifiers (Anthropic's Constitutional Classifiers is cited), linear probes trained on internal activations, static analysis tools (CodeShield is cited), and manual and automated abuse monitoring. Limitations recorded include computational cost and latency at higher input volumes, circumvention by decomposing malicious requests into benign-appearing substeps or distributing queries across multiple accounts, the need to retrain probes for each model version, static tools missing context-specific issues, and the difficulty of defining abusive patterns without violating user privacy or flagging benign research behavior.
  • Access control — staged deployment, beginning with vetted cybersecurity professionals or research partners under monitoring agreements and broadening as controls are validated. The report notes a tradeoff: stringent vetting requirements may exclude smaller organizations that lack the resources or teams to meet them, and those organizations may lose access to beneficial AI.
  • Ecosystem mitigations — secure information-sharing networks (the FMF's own information-sharing agreement is cited), defensive systems and research, and programs for cyberdefense including OpenAI's trusted-access programs and Anthropic's critical-infrastructure defense partnerships, the latter described as enabling defensive applications such as emulating cyber attacks on water treatment plants. Recorded limitations include the difficulty of building trusted channels across public–private and international boundaries, the legal, privacy and commercial-sensitivity constraints on sharing, the concentration of advanced defensive capability in well-resourced organizations, and the difficulty of verifying who qualifies as a trusted cybersecurity user given fabricated credentials, changing affiliations and insider threats.

The report also records that open-source models, artifacts and defensive tooling can act as a catalyst for exploration in the security ecosystem, and that some developers publish open-source tools and measurement suites for identifying malicious activity, extracting insights from threat-intelligence reports, and examining automated patching of vulnerable systems.

Mitigation assessments

The report states that after implementing safeguards, developers must verify that they reduce risk to acceptable levels under realistic operating conditions, and distinguishes testing individual safeguards (for example, classifier precision and recall, or behavioral guardrails against prompt variations) from system-level testing of all safeguards working together, which it describes as giving a more realistic picture and able to reveal emergent vulnerabilities that do not appear when components are tested in isolation.

Three approaches are identified. Empirical testing by internal safety teams, benefiting from institutional knowledge but subject to organizational blind spots, using prompts the model should refuse entirely, queries requiring cautious or limited responses, and prompts where full compliance is acceptable, tested repeatedly, in combination with jailbreaks and across multi-turn interactions; third-party assessors, who may use publicly available benchmarks or privately developed harmful questions, are described as providing additional coverage and a different depth of domain evaluation. Red teaming requiring both red-team expertise and specialized domain knowledge, conducted on systems as they will be deployed, with external cybersecurity firms bringing established testing frameworks and cross-organizational experience and serving firms that lack internal expertise. Ongoing deployment monitoring, on the argument that mitigation assessments retain usefulness longer than capability benchmarks because they measure safeguard robustness rather than raw capability, which saturates — while noting that safeguard performance can still decline as attackers develop new circumvention techniques, so detection-system accuracy rates, attack-complexity trends and intervention success rates should be tracked.

Continuing work

The report's closing section identifies three areas of ongoing work. Cyber capability assessment and threat modeling: developing automated evaluations and training metrics that identify concerning capabilities earlier in development, addressing models that might conceal their capabilities during evaluation, and collaborating on standardized evaluation frameworks, shared red-teaming methodologies and common benchmarks; it notes that advances in interpretability and transparency may give deeper insight into how models acquire and apply cyber knowledge. Mitigation design and validation: safety mechanisms that resist jailbreaking and prompt injection, techniques to limit models' ability to reason about or execute offensive operations, standardized approaches for testing adversarial robustness, and empirical validation against realistic attack scenarios. Risk–benefit tradeoffs: the report states that developers can provide transparency about their cyber capability thresholds and decision processes, but that the ultimate determination of acceptable risk–benefit tradeoffs requires input from cybersecurity experts, policymakers, diverse stakeholders and appropriate governance structures.

Relationships