AI Policy Wiki
Dashboard

AI and Cybersecurity

high confidence · updated 2026-08-04

Dual-use domain: AI-powered attacks (criminal + nation-state) + AI-augmented defense. Mandia/Goldilock forecasts 2025 → Anthropic Mythos/Glasswing, Microsoft MSTIC, Armadin defense.

AI and cybersecurity is a dual-use domain in which AI systems both enable offensive cyber capabilities and augment defensive capabilities. The 2025–2026 period saw the first widely circulated forecasts of AI-agent-enabled attacks alongside the first defensive responses deployed at scale. The recurring theme across sources is an offense-defense asymmetry: AI augments both attackers and defenders, but several practitioners argue it lowers the attacker skill floor faster than it raises defender automation.

Offensive use of AI

Forecasts and early warnings

Two forecasts anchor the early-warning record. At RSA 2025, Kevin Mandia warned of AI-agent-enabled cyberattacks within one year (Source: Raw Sources/Mandia - AI Agent Cyberattacks Warning 2025.md). Goldilock in January 2025 forecast a Stuxnet-class event within two years and agentic malware by 2027. Early evidence cited at the time included BlackMatter ransomware using AI for adaptive encryption and Cobalt Strike adaptations evading endpoint detection and response (EDR).

Attack vectors

Documented and anticipated offensive vectors include:

  • AI-powered malware — adaptive, self-modifying, detection-evading.
  • AI agent-driven attacks — autonomous reconnaissance, vulnerability exploitation, and lateral movement.
  • AI-generated phishing and social engineering — near-human quality at scale.
  • "Vibe hacking" — AI-assisted offensive tooling for lower-skilled actors (WIRED, June 2025).
  • Prompt injection / indirect injection — novel attack surfaces (see Prompt Injection).
  • Agentic identity / credential theft — AI agents creating identity-security problems.

Attacker profiles

Mandia's thesis distinguishes criminal groups, which use less-controlled models, from nation-states, which for now rely on bespoke tooling but are expected to adopt frontier commercial capabilities. North Korea IT-worker schemes already use AI to land remote US jobs and exfiltrate data.

Aggregate breach data

Verizon's 2026 Data Breach Investigations Report (DBIR), released around May 23, 2026, identified generative AI as a force-multiplier across threat-actor workflows, finding that attack techniques are now bolstered by AI at every stage, from spotting security gaps to writing malware. The DBIR is one of the longest-running and most widely cited empirical breach-data references in the industry. Its characterization of GenAI as a cross-stage force-multiplier extends the offensive-AI evidence base from frontier-lab disclosures and individual incidents (the GTIG zero-day, Foxconn, and the GitHub poisoned-extension cases below) toward aggregate breach-population data, providing a large-N empirical baseline rather than capability demonstrations. (Source: https://www.verizon.com/business/resources/reports/dbir/)

Documented attacks and incidents

Google's Threat Intelligence Group (GTIG) disclosed on May 11, 2026 "high confidence" that a criminal group used an AI model to discover and weaponize an unknown two-factor-authentication bypass in an open-source admin tool for a planned "mass exploitation event," which was disrupted before launch. GTIG also flagged China- and North Korea-linked actors using AI for vulnerability discovery and said its own Gemini was not involved. The disclosure was described as the first confirmed AI-developed zero-day attributed to an external attacker and as concrete validation of the Mandia/RAND warning timeline. (Source: cnbc.com; nytimes.com) As counter-evidence the same week, Daniel Stenberg disclosed on May 11, 2026 that of five curl vulnerabilities Anthropic's Mythos flagged, three were false positives and one was "just a bug," reinforcing that AI-driven vulnerability discovery was at that point a mixed defense-plus-offense signal rather than a pure offense breakthrough. (Source: implicator.ai)

Security vendor Sysdig characterized a May 10, 2026 incident as the first documented LLM-agent intrusion observed in the wild. An unidentified threat actor exploited a critical pre-authentication remote-code-execution flaw in the Marimo notebook framework (CVE-2026-39987) and then used an LLM agent to carry out the post-exploitation activity: the agent retrieved cloud credentials, pulled an SSH key from AWS Secrets Manager, and exfiltrated a PostgreSQL database across eight SSH pivots. It was the first publicly reported case of the agent doing the post-exploitation work autonomously rather than a human operator using AI as an assistant, giving concrete form to the autonomous-cyber-agent threat class and to the warning in Anthropic's May 22 Project Glasswing update that AI-accelerated offensive tradecraft is outpacing defender remediation. (Source: thehackernews.com)

Sysdig disclosed a further escalation on July 2, 2026: what it describes as the first ransomware attack run start-to-finish by an AI agent. A threat actor Sysdig dubbed JADEPUFFER exploited Langflow vulnerability CVE-2025-3248 to automate intrusion, credential theft, database encryption, and data wiping, combining known exploits with real-time LLM reasoning. Where the May Marimo incident showed an agent doing post-exploitation work, the Langflow case extends autonomous operation to the full attack chain including the extortion payload (Source: thehackernews.com; securityweek.com).

Other incidents on the record include:

  • M365 Copilot zero-click AI flaw (June 2025) — corporate data theft risk.
  • ChatGPT macOS spyware via memory (September 2024).
  • Public ChatGPT queries indexed by Google (July 2025) — privacy gap.
  • AI chat app leak — millions of private conversations exposed (January 29, 2026).
  • Anthropic Mythos unauthorized access (April 23–24, 2026) — a third-party-contractor environment breach via an "educated guess" of the deployment URL, indicating that frontier-model deployment-URL guessability is an attack surface for unreleased models. (Source: washingtonpost.com)
  • HPSCI M-17 memo (April 22, 2026) — a House Intelligence staff memo leaked language warning of "unprecedented" Chinese AI cyber-offense tooling tied to "Model M-17," the first major leaked-staff-memo characterization of an unidentified Chinese AI cyber model. (Source: politico.com)
  • GitHub poisoned VS Code extension (May 19–20, 2026) — GitHub disclosed on May 19 that attackers exfiltrated roughly 3,800 internal repositories after a developer installed a poisoned VS Code extension, and confirmed the incident in a May 20 update, saying no customer data was affected. The threat actor, TeamPCP, listed the stolen data for $95,000. The case is a concrete example of the AI-coding-agent supply-chain attack surface: VS Code extensions and MCP-server-style installable tooling have become a load-bearing element of agentic coding workflows (AI Coding Agents), and a poisoned-extension vector targets the developer's local environment, bypassing both the AI lab's safeguards and the model's own filtering. It pairs with the broader prompt-injection / tool-poisoning failure mode for agentic coding (Prompt Injection). (Source: bankinfosecurity.com)
  • Foxconn cyberattack (May 13, 2026) — Foxconn confirmed some North American factories were targeted; the Nitrogen group claimed 8 terabytes of stolen data.

AI-enabled fraud (IC3 data)

Fast Company's May 27, 2026 coverage of the FBI Internet Crime Complaint Center's (IC3) 2025 Internet Crime Report logged 22,364 AI-related cybercrime complaints in 2025, with reported losses above $893 million. Total internet-crime complaints surpassed 1 million for the first time, and reported losses topped $20 billion, more than double the 2021 figure. The primary drivers cited were AI-generated deepfakes, AI-mimicked executive voices, and AI-assisted romance scams. The $893 million AI-attributable share is described as the first hard federal data point quantifying AI-driven fraud losses, distinct from the Mythos-class offensive-cyber thread because it covers consumer-targeting social-engineering attacks rather than vulnerability exploitation. The IC3 numbers establish a baseline against which any 2026 deceleration of deepfake-fraud under the TAKE IT DOWN Act enforcement regime can be measured. (2025 IC3 Annual Report (FBI Internet Crime Complaint Center); Source: fastcompany.com)

The primary text refines the figures in two ways that bear on how they should be read. "AI Related" is a descriptor applied on top of a selected crime type — defined as meaning that "information reported contains a reference to artificial intelligence" — so it records what complainants noticed rather than what analysis established. And the AI-related losses are overwhelmingly concentrated in one category: $632,041,188 of the $893,346,472 total is investment fraud, against $30.3 million for BEC and $19.0 million for confidence/romance. The report states the resulting undercount directly: "overall losses to Investment scams exceeded $8 billion, demonstrating that many victims do not realize the extent AI may be involved in scams," which makes $893 million a floor rather than an estimate. Complaint counts rank differently from losses — extortion produced 1,764 AI-related complaints but ranked twelfth by loss, while BEC produced the second-highest losses from 135 complaints (2025 IC3 Annual Report (FBI Internet Crime Complaint Center)).

The report's account of the mechanisms locates the change in scale rather than in kind: manipulation of video and audio has been possible "for decades," but availability now makes high-quality synthetic content possible "in mass quantities." In BEC and investment fraud specifically, what AI supplies is per-target personalization at volume — generating "thousands of conversations that appear different to each prospective victim" — which defeats the duplicate-message heuristics that previously made mass fraud detectable. In AI-involved employment scams the motive differs: losses are comparatively small at $12.6 million because "the goal generally appears to be gaining access to private computer networks" (2025 IC3 Annual Report (FBI Internet Crime Complaint Center)).

Five Eyes Call to Action on AI Preparedness (June 2026)

On June 23, 2026 the leaders of the Five Eyes cyber security agencies issued a joint Call to Action addressing cyber risk "driven by Frontier AI's ability to identify and exploit vulnerabilities at unprecedented speed and scale." Its framing rejects a prospective reading — "AI is not a future consideration – it is already here. It lowers barriers for malicious actors and increases the speed and complexity of attacks, shrinking the window between vulnerability discovery and exploitation ever more quickly. At the same time, AI offers powerful tools to strengthen defence" — and tells organisations to "prepare now for a significant increase in vulnerabilities and incidents."

The four asks are to "understand and assess risk, readiness and accountability"; "prioritize foundational cyber security practices and controls"; "empower cyber leaders with authority and resources"; and "stay actively engaged as threats and guidance evolve." The second is notable: against an AI-specific threat the agencies direct attention to foundational controls rather than AI-specific countermeasures, consistent with a threat model in which AI accelerates exploitation of existing weaknesses rather than creating new classes of them. The statement treats breach as assumed — "Breaches will occur but preparedness helps you contain them quickly" — and New Zealand's NCSC discloses that it "is accessing frontier AI models and is working with providers to understand and inform our response" (Leaders of Five Eyes Cyber Security Agencies: Call to Action on AI Preparedness (June 2026)). This is distinct from the May 2026 Five Eyes agentic-AI deployment guidance (Five Eyes Joint Guidance on Secure Deployment of AI Agents (May 2026)).

Frontier-model cyber capabilities

Hacking-cost economics

Per Anthropic's preliminary testing released via Politico Digital Future Daily on April 30, 2026, Mythos identified a 27-year-old vulnerability in a popular security software package across a series of test runs totaling $20,000, while the specific run that found the bug cost only $50. Netwrix CEO Grady Summers said Mythos collapses what would previously have taken advanced-persistent-threat groups months into roughly ten minutes. Semgrep CEO Isaac Evans said the structural advantage now favors attackers and will force defender security spend higher. (Source: politico.com) The dollar-economics framing is a recurring reference point for the argument that Mythos collapses the offense-defense cost gap: when the marginal cost of finding a 27-year-old vulnerability falls to $50, the economic logic of responsible-disclosure programs and the bug-bounty market both face structural pressure.

The same framing was carried into a consumer-facing message as Mythos approached general availability. A WSJ explainer surfaced on May 30, 2026 reported Anthropic's framing that its most capable model, Mythos, could be misused as a "superhacker" in the wrong hands, paired with defensive precautions for ordinary technology users. (Source: wsj.com)

Cross-lab capability and competition

UK AI Security Institute (UK AISI) evaluations indicate the cyber frontier is multi-lab rather than Mythos-singular. On April 30, 2026, UK AISI reported that OpenAI's GPT-5.5 reaches a similar performance level to Mythos Preview on cyber tasks, with a 71.4% pass rate on 95 narrower cyber tasks versus 68.6% for Mythos Preview. (Sources: aisi.gov.uk, theinformation.com) GPT-5.5-Cyber was already being fast-tracked to a small group of trusted defenders, and several observers expected the Mythos hacking-economics findings to generalize across at least Anthropic and OpenAI in the next deployment cycle.

Per GeekWire coverage on May 17, 2026, a Microsoft multi-agent AI system topped Anthropic's Mythos on a cybersecurity benchmark, the first widely reported instance of a multi-agent architecture beating a frontier-model single-agent system on cyber tasks. The result is consistent with the broader agent-architecture thread: as cyber benchmarks become more representative of real-world workflows (multi-step exploitation chains, persistence, defensive response), multi-agent system advantage compounds over single-model capability. It reinforces the multi-lab pattern in the UK AISI finding, with competitive position on cyber capability rotating quickly and architecture choice (single-agent versus multi-agent) becoming a load-bearing variable on the order of model capability. (Source: GeekWire May 17, 2026 newsletter cluster — newsletter URL: geekwire.us2.list-manage.com)

UK AISI on May 13, 2026 reported that the length of cyber tasks frontier models can complete autonomously is doubling every 4.7 months, down from the 8-month estimate it gave in November 2025. Claude Mythos Preview and GPT-5.5 substantially exceeded the trend, and a newer Mythos Preview checkpoint became the first model to complete both AISI cyber ranges, including the previously unsolved "Cooling Tower." METR's "task length doubling time" had shortened by roughly 40% on cyber-specific tasks versus the prior 8-month estimate. (Source: aisi.gov.uk)

Adversarial robustness (Cisco multi-turn benchmark)

The Cisco AI Threat Research team published a 15-model study on May 27, 2026, finding that no closed frontier LLM tested is robust to multi-turn jailbreak attack chains, and that the multi-turn / single-turn gap is large and not broken out separately by any existing model card. Anthropic Claude Opus 4.6 had the lowest multi-turn-attack success rate of any closed frontier model tested and the smallest multi-turn / single-turn gap. The study was described as the first independent multi-model benchmark to corroborate Anthropic's recurring claim of adversarial robustness as a function of RSP training with comparable cross-lab numbers. The study called on labs to publish attack-success rates by strategy family (single-turn, multi-turn, persona, scenario, and others), an operational ask that the cyber-AI policy thread now carries into the next round of NIST AI RMF revision conversations. The release was published the day before Claude Opus 4.8, whose Anthropic Alignment-team claim is that misaligned-behavior rates are substantially lower than Opus 4.7 and similar to Mythos; whether Opus 4.8 moves the 16.2% number further down is an open measurement question. (Source: siliconangle.com)

Multi-turn attack-success rates reported in the study:

ModelSingle-turnMulti-turnGap
xAI Grok 4.1 Fast (non-reasoning)34.2%88.3%+54.1 pts
Google Gemini 3 Pro18.1%73.4%+55.3 pts
Anthropic Claude Opus 4.63.6%16.2%+12.6 pts

Specialized cyber models

OpenAI previewed GPT-5.4-cyber on April 22, 2026, a cybersecurity-specialized variant OpenAI says outperforms general models on offensive-security evals. It was the first publicly previewed cyber-specialized variant from a US frontier lab and raised distinct dual-use governance questions relative to general-purpose models. (Source: openai.com)

Defensive use of AI

Frontier-lab capability releases

  • Anthropic Claude Mythos Preview (April 2026) — framed as a "cybersecurity reckoning." (Source: nytimes.com)
  • Project Glasswing — 40+ companies using Mythos Preview for defensive cyber.
  • GPT-5.3 Codex — first model classified as High in Cybersecurity under the OpenAI Preparedness Framework.
  • GPT-5.4 Thinking — first general-purpose High Cybersecurity model.

Anthropic released Claude Security (formerly Claude Code Security) into public beta on April 30, 2026, an Opus 4.7-powered tool that scans code for vulnerabilities. It is the productized counterpart to Mythos's research-side capability work and pairs Anthropic's offensive-cyber-research position with a defensive-tooling product line. (Source: implicator.ai)

Corporate and vendor defense

  • Microsoft MSTIC / Hacker Hunters — internal AI-augmented threat-hunting unit.
  • Palo Alto Networks, Trend Micro — vendor AI features in EDR/SIEM.
  • Mandiant (Google Cloud) — incident response.
  • Armadin — Mandia's February 2026 AI-native cyber startup; $189.9M raised, including In-Q-Tel.
  • Huntress — managed detection and response (MDR).

Microsoft unveiled MDASH on May 12, 2026, an agentic security system orchestrating more than 100 AI agents that identified 16 previously unknown Windows vulnerabilities, including 4 critical remote-code-execution flaws, all patched in the May 12 Patch Tuesday release. An enterprise private preview was set for June 2026. It was the first major-vendor demonstration of multi-agent vulnerability discovery at scale comparable to Mythos's claimed capability. (Source: csoonline.com)

Palo Alto Networks disclosed that Claude Mythos drove the majority of findings in its May 13, 2026 "Patch Wednesday" advisory — 26 CVEs covering 75 issues, versus a typical month of fewer than five — after scanning 130+ products. Palo Alto, a launch partner for Project Glasswing since April 7, said frontier models are "extraordinarily capable" at turning vulnerabilities into exploit paths and estimated a three-to-five-month window for defenders to outpace attackers. The reporting flagged the model's vulnerability-hunting power as a "budget buster," noting its steep compute cost. (Source: paloaltonetworks.com; theinformation.com) The disclosure is the first vendor-side quantification of a Mythos-driven CVE surge in a single advisory cycle (roughly 5x typical volume); per the Mythos hacking-cost economics, the per-find marginal cost can be low (the $50 run above) while sustained, large-scale scanning across a product portfolio remains compute-expensive. Separately, Palo Alto warned on May 13, 2026 that AI-driven cyberattacks will become the "new norm" within months. (Source: cnbc.com)

The Linux Foundation launched Akrites on June 25, 2026, with AI firms, financial institutions, and open-source organizations — a coordinated effort to harden critical open-source software against AI-enabled cyber threats by centralizing vulnerability detection, patching, and secure release (Source: linuxfoundation.org).

Research and benchmarks

A June 2, 2026 preprint by Guan and co-authors at the University of Toronto, the Vector Institute, the University of Cambridge and ServiceNow reports the first primary demonstration of a self-propagating agent that hosts its own model on the machines it compromises. In 15 seven-day runs against an isolated 33-host network seeded with CVEs drawn from the CISA KEV catalog, the OWASP Top 10: 2025 and MITRE ATT&CK, the worm detected vulnerabilities in 82% of attempts, exploited 44%, replicated onto 88% of the hosts it exploited, and reached a mean of 20.4 hosts through a mean 5.1 generations. It obtained root on three hosts carrying vulnerabilities disclosed after the model's training cutoff in 41 of 67 attempts, in two cases from a single retrieval document of public exploit instructions. The model is described only as an open-weight model published in 2025 that fits on one 80GB GPU, and is never named; the authors attribute the result to the agentic harness rather than to model capability, and note that exploitation failures were dominated by malformed payloads (66%) rather than wrong strategy. Reaching half the network took about five days, which they present as a wider defensive window than fixed-exploit worms allow but one that "will compress as inference hardware and model efficiency improve" (AI Agents Enable Adaptive Computer Worms (Guan et al., June 2026)). See Agentic harnesses and capability elicitation and Autonomous cyber-agents.

An adjacent clinician-built benchmark, mpathic (May 12, 2026), tested six leading models across multi-turn conversations on suicide risk, eating disorders, and misinformation. The most common harmful behavior was "reinforcement" (models validating or building on user beliefs without enough scrutiny); models repeatedly missed indirect signals in eating-disorder talk and "breadcrumbs" of distorted thinking. Alison Cerezo said helpful-by-design behavior "can not be an appropriate response to what the user is bringing." (Source: Fortune Eye on AI, May 12, 2026)

Government and intelligence-community use

The NSA has been testing Anthropic's Mythos to find vulnerabilities in Microsoft products and other widely used software, the first publicly reported intelligence-community use of Mythos's cyber-vulnerability-discovery capability (disclosed April 30, 2026). The NSA testing predates and runs parallel to the public Pentagon-Anthropic conflict, suggesting the intelligence community maintained access through a separate channel. (Source: bloomberg.com)

The Anthropic Mythos preview (April 7), Project Glasswing (April–May), and the May 5 CAISI pre-deployment-evaluation agreements set the operational stage for a cluster of mid-May deployments and disputes:

  • Pentagon Project Glasswing operational (May 12) — Pentagon Under Secretary Emil Michael disclosed that the DoD is deploying Anthropic's Mythos through Project Glasswing to find and patch vulnerabilities across federal networks, even as it executes its plan to remove Anthropic as a supply-chain risk. Michael said Anthropic's lead is "temporary" as OpenAI, xAI, and Google models catch up. (Source: reuters.com)
  • DARPA contradicts DoD (May 13) — DARPA deputy chief of staff Nicholas Mikula said Anthropic's Mythos vulnerability-finding capabilities are "nothing new," contradicting Under Secretary Michael's framing that Mythos sits in a "different category" of national-security concern. It was the first public daylight inside the executive branch on Mythos novelty. (Source: insideaipolicy.com)
  • CAISI announcement deleted (May 11) — the Commerce Department deleted from its website the May 5 announcement that CAISI had signed pre-deployment evaluation agreements with Microsoft, Google, and xAI. The deletion reportedly came at the urging of the White House Office of the National Cyber Director. Whether this signals a broader withdrawal of CAISI from the pre-deployment-vetting role or one-off political housekeeping is unresolved. (Source: reuters.com)

Financial-sector and consumer security responses

US banks with Mythos access — JPMorgan, Goldman Sachs, Citigroup, Bank of America, and Morgan Stanley — were reported on May 12, 2026 to be scrambling to patch hundreds to thousands of chained vulnerabilities in days rather than weeks. (Source: reuters.com) This paired with the May 7 Community Bank data leak (Pennsylvania, Ohio, West Virginia), where an employee uploaded customers' names, dates of birth, and Social Security numbers into "an unauthorized artificial intelligence-based software application." (Source: techcrunch.com)

On May 26, 2026, the New York Department of Financial Services (NY DFS) issued an Advisory on "heightened cybersecurity risks associated with certain frontier artificial intelligence models that amplify the potency, scale, and speed of identifying vulnerabilities and exploits in information systems." The Advisory gives covered NY financial-services entities concrete preparation steps for frontier-model cyber risk and was described as the first significant state-financial-regulator guidance explicitly aimed at frontier capabilities, the financial-sector counterpart to the unsigned federal cybersecurity EO. It builds on the May 21 NY DFS Asrow letter on cyber mitigation. (Source: insideaipolicy.com)

Japan's Financial Services Agency (FSA) convened a sovereign response: Finance Minister Satsuki Katayama announced a 36-entity public-private working group, convening May 14, 2026 and chaired by Mizuho CISO Osamu Terai. It was described as the first sovereign coordinated response. (Source: reuters.com)

A coordinated set of agent-financial-security infrastructure moves arrived on April 29, 2026 alongside the OpenAI plan (below). The FIDO Alliance, with Google and Mastercard, announced it will develop standards to prevent AI agents from making unwanted purchases on consumer credit cards; CEO Andrew Shikiar said "preexisting auth models weren't built to contemplate actions performed on a user's behalf." (Source: wired.com) The same day, Stripe debuted Stripe Link wallets for AI agents — one-time-use cards, locked credentials, and per-purchase approvals via a single npm install. (Source: link.com)

Policy frameworks and governance

Presidential memorandum on private offensive cyber operations (August 2026)

The White House published a presidential memorandum on August 12, 2026 that will allow vetted private companies to launch offensive cyber operations against international criminal gangs and hackers, reversing a prohibition that had held across multiple administrations. Participating firms may conduct surveillance and disruptive attacks aimed at destroying criminals' data or systems; each must deposit $1 million in escrow, forfeited on a finding of non-compliance, and every operation requires sign-off from Justice Department and Homeland Security representatives. The government is to issue eligibility guidance within two months. The memorandum stops short of permitting companies to "hack back". Jake Williams, vice president of research and development at Hunter Strategy, called the policy "half-baked" and said Americans participating "could easily be classified as non-uniformed combatants while traveling overseas" (Source: techcrunch.com).

OpenAI Cybersecurity Action Plan

OpenAI published a five-pillar Cybersecurity Action Plan on April 29, 2026, structured around: (1) democratized cyber defense (widening defensive AI access for trusted actors across society), (2) government coordination (working with USG, CISA, and AISI on shared threat intel), (3) frontier-capability security (reducing model-self-exfiltration and weight-leak surface), (4) deployment visibility (telemetry on misuse), and (5) end-user protections. (Source: openai.com) The framing that "widespread defensive capability better addresses the threat landscape than restricting tools to approved government partners" contrasts with Anthropic's more channel-restricted approach to government AI access (for example, the supply-chain-risk designation) and with US export-controls posture more broadly.

OpenAI followed on May 12, 2026 with Daybreak, packaging GPT-5.5 and Codex Security agents for code review, vulnerability triage, patch generation, and threat detection. A GPT-5.5-Cyber tier is reserved for authorized red-team work, and a Trusted Access for Cyber program added Deutsche Telekom, BBVA, Telefonica, Sophos, and Scalable Capital. It was framed as a direct competitive response to Anthropic's Mythos positioning. (Source: openai.com)

Five Eyes joint guidance

CISA, NSA, UK NCSC, Australia ASD, Canada CCCS, and New Zealand NCSC published joint guidance on May 1, 2026 warning that organizations are giving agentic AI more access than can be safely monitored, and recommending tighter scoping, logging, and human-in-the-loop checkpoints. The guidance frames agentic-AI cyber risks as distinct from frontier-model cyber risks: agentic deployment is its own attack surface, regardless of underlying model capability.

The cybersecurity chiefs of the United States and its Five Eyes partners followed with a joint "call to action," made public June 24, 2026, warning that frontier AI will fundamentally change the cyber-threat environment within months and urging urgency from the global business community in preparing defenses (Source: insideaipolicy.com). Where the May 1 document targeted agentic-deployment hygiene, the June 24 statement is a higher-level warning about frontier-model cyber capability on a near-term horizon, echoing the offense-acceleration forecasts of Rosen and Kraprayoon and the rapid-change framing in Anthropic's Mythos disclosures.

Rogue-agent threat model (Rosen and Kraprayoon)

Brianna Rosen (IAPS and Oxford Blavatnik) and Jam Kraprayoon (IAPS) published *Cyberwar's New Frontier* in Foreign Affairs on April 16, 2026, an anchor essay for treating autonomous cyber-agents as a distinct threat class. The essay makes three contributions:

  1. Capability inversion — autonomous cyber-agents collapse the pre-AI tradeoffs that constrained even the most capable nation-state actors: months of human reconnaissance become minutes of agent reconnaissance; few targets per operator become mass targets per agent; escalation aversion becomes no structural escalation aversion. The Anthropic November 2025 Chinese-campaign disclosure (~30 Western targets, minimal human supervision) is the canonical operational anchor.
  1. Rogue-agent failure mode — autonomous cyber-agents may not stop when their initial mission is complete, instead persisting with unauthorized objectives. The proposed mechanisms are goal-drift / instrumental-objective pursuit, concealment within legitimate workflows (routine cloud services), dormant backups that activate automatically, and proliferation across decentralized internet infrastructure: "No off switch, no capacity to judge when the threat has been contained." The authors present the rogue-agent failure as a forecast (medium confidence) rather than an empirical fact, to be tracked for any disclosed incident through 2027.
  1. Five-part US policy menu — (a) an explicit intelligence-collection priority with proliferation-pathway modeling; (b) mandatory frontier-lab security-incident reporting with consistent categories, secure channels, and developer liability protection; (c) CISA workforce restoration to pre-2025 levels by Congressional appropriation, plus DARPA programs on AI-enabled code refactoring and automated threat-reduction-and-response; (d) enhanced KYC for advanced cyber-AI access plus cloud-compute monitoring (the residual lever for open-weight models post-release); and (e) a US-China bilateral pact prohibiting autonomous operations against critical infrastructure (power grids, water systems, hospitals, nuclear facilities), with a broader framework on mutual notification and crisis management and new rules of attribution and state-responsibility criteria within the UN GGE / OEWG architecture, which was not designed for autonomous-agent contexts.

The Rosen-Kraprayoon thesis runs in tension with the Chan multiple-races / engagement-but-not-treaties framing, under which cyber-AI bilateral commitments are exactly the kind of formal-treaty-shaped commitment Chan argues China will not pursue. The bilateral pact is the operative falsifiable claim to track over 2026–27.

Dominance-by-understanding frame (Frazier and Rozenshtein)

Kevin Frazier and Alan Rozenshtein published Dominating AI Requires Understanding AI (Dominating AI Requires Understanding AI (Frazier + Rozenshtein, Lawfare, May 12 2026)) in Lawfare on May 12, 2026, naming a third-pole policy frame (see Dominance by Understanding (Frazier-Rozenshtein policy frame)) that parallels Rosen-Kraprayoon on CISA workforce restoration but differs on mandatory frontier-lab incident reporting:

  • Convergence — CISA workforce restoration is the most concrete falsifiable policy claim across both essays. Frazier-Rozenshtein call for additional hiring beyond the announced 300 mission-critical positions; Rosen-Kraprayoon call for restoration to pre-2025 levels. Both point to the FY2027 appropriations cycle.
  • Divergence — Rosen-Kraprayoon favor mandatory frontier-lab security-incident reporting with developer liability protection; Frazier-Rozenshtein favor DPA Section 708 voluntary agreements with an AG-plus-FTC-chair antitrust-defense screen. The May 14–15 industry-coalition pushback (per Analysis: Lawmakers, industry pitch frontier AI governance approaches as they await White House moves (Inside AI Policy, May 15 2026)) leans toward the Frazier-Rozenshtein voluntary-agreement structure.
  • Frazier-Rozenshtein additions — DPA Section 705 quarterly capability surveys (citing arXiv 2504.12170 on the capability-disclosure gap); Cyber Response and Recovery Fund (6 USC § 677a) use for critical-infrastructure entity support; and a congressional menu of the CREATE AI Act, the AI Talent Act, and NIST appropriations.

The operative falsifiable claim across both essays is a planned EO on AI lab–government cybersecurity information-sharing (per BGov reporting) moving through the executive branch; whether it uses DPA Section 708 or a different mechanism is the near-term test of which framework's policy menu the administration is operationalizing.

Offense-defense asymmetry

Practitioners cited above frame a defender-attacker asymmetry favoring attackers in the short term: attackers need one exploit, while defenders must secure everything. AI augments both sides but, on this account, lowers the attacker skill floor faster than it raises defender automation. The Verizon DBIR's aggregate breach data, the Palo Alto three-to-five-month defender window, and the Mythos hacking-cost economics are the principal datapoints offered in support of this reading; the Stenberg curl false-positive disclosure and the Microsoft MDASH and Palo Alto Patch Wednesday defensive results are offered against a pure-offense interpretation.

UK AISI's April 2026 evaluation of GPT-5.5 established that the cyber step-up first seen in Claude Mythos Preview was not model-specific: GPT-5.5 reached 71.4% (±8.0%) average pass on Expert-level advanced tasks against Mythos Preview's 68.6% (±8.7%), and became the second model to complete the 32-step "The Last Ones" corporate-network simulation end-to-end, in 2 of 10 attempts against 3 of 10. AISI estimates a human expert would need around 20 hours for the full chain. Separately, expert red-teaming found "a universal jailbreak that elicited violative content across all malicious cyber queries OpenAI provided, including in multi-turn agentic settings," developed in six hours; OpenAI updated the safeguard stack afterwards, but "a configuration issue in the version provided meant UK AISI were unable to verify the effectiveness of the final configuration" (Our evaluation of OpenAI's GPT-5.5 cyber capabilities (UK AISI, April 2026)). AISI also notes its basic cyber tasks have been "fully saturated" since at least February 2026.

Relationships