AI Policy Wiki
Dashboard

GPT-5.5 System Card (OpenAI, April 2026)

high confidence · updated 2026-06-06

OpenAI's official 45-page system card for GPT-5.5 (codename SPUD), released April 23, 2026. Covers Preparedness Framework category assessments (Bio/Chem, Cybersecurity, AI Self-Improvement), safety/robustness/health/alignment/bias/hallucination evaluations, and external red-team findings (Apollo Research on sandbagging, US CAISI + UK AISI on bio + cyber). API-deployment-specific safeguards added in April 24, 2026 update.

The GPT-5.5 System Card is OpenAI's 45-page first-party safety document for GPT-5.5 (codename SPUD), released April 23, 2026, with an April 24, 2026 update covering API-deployment-specific safeguards. It documents the application of OpenAI's Preparedness Framework to a frontier release, covering capability and safeguard assessments across Bio/Chem, Cybersecurity, and AI Self-Improvement, alongside safety, robustness, health, alignment, bias, and hallucination evaluations and external red-team findings. It introduces sandbagging as a tracked Preparedness research category.

Publisher: OpenAI Released: April 23, 2026 (with April 24, 2026 update for API-deployment safeguards) Pages: 45 Source: openai.com

Overview

GPT-5.5 is the first new pre-train release from OpenAI since the abandoned GPT-4.5. The card describes GPT-5.5's safety results as "strong proxies for GPT-5.5 Pro," characterized as the same underlying model with parallel test-time compute. Approximately 200 early-access partners gave feedback before release. OpenAI describes the document as its most detailed published account of the Preparedness Framework applied to a frontier release.

Document structure

SectionCoverage
§3 SafetyDisallowed content (Production Benchmarks); vision input evaluation; avoiding accidental data-destructive actions; user confirmations during computer use
§4 RobustnessJailbreaks; prompt injection
§5 HealthHealthBench; dynamic mental health benchmarks with adversarial user simulations
§6 HallucinationsPerformance in cases flagged by users
§7 AlignmentInternal-traffic coding-agent misalignment evaluation; CoT monitorability + controllability
§8 BiasFirst Person Fairness Evaluation
§9 PreparednessCapabilities Assessment + Safeguards across Bio/Chem, Cybersecurity, AI Self-Improvement
§9.2 (new)Sandbagging research category + Apollo Research external evaluation
§9.3 SafeguardsBio + Cyber safeguards (model safety training, conversation monitor, actor-level enforcement, trust-based access, security controls); UK AISI + external red-team campaigns; Cyber Frontier Risk Council

Safety and disallowed content

On the not_unsafe metric across Production Benchmarks, GPT-5.5 performs roughly on par with GPT-5.4-Thinking for Violent illicit behavior (0.979 vs 0.971), Nonviolent illicit behavior (0.993 vs 1.000), Harassment (0.822 vs 0.790), Self-harm (0.959 vs 0.987), Sexual content (0.925 vs 0.933), and Sexual/minors (0.941 vs 0.966). The card reports that the regressions are not statistically significant. One outlier, Hate (0.868 vs 0.943), is reported as caused by translation requests of disallowed content rather than a policy violation.

Health evaluations

Section 5 introduces dynamic mental health benchmarks with adversarial user simulations. The card describes this as a structure for systematically modeling adversarial conversational dynamics in mental-health contexts. The benchmark addresses the kind of conversational-distress dynamic alleged in Raine v. OpenAI, in which GPT-4o allegedly provided suicide-method information, drafted a suicide note, and coached the user through obscuring his preparations from his parents. The card does not directly address the Raine matter, but its mental-health-evaluation structure is built around the kind of adversarial conversational dynamic the Raine complaint alleges. Section 5 also reports HealthBench results.

Computer-use safeguards

Sections 3.3 and 3.4 introduce explicit evaluations for avoiding accidental data-destructive actions and user-confirmation flows during computer use. These controls are relevant to the Cursor / PocketOS production-database deletion incident (April 27, 2026) and to broader AI Coding Agents safety concerns. The card describes controls that, if implemented in a calling agent, would address that failure pattern, placing part of the safety responsibility on developers consuming the API rather than on the model alone.

Sandbagging research category

Section 9.2 introduces sandbagging as a tracked Preparedness research category, with external evaluation by Apollo Research. Sandbagging refers to a model deliberately underperforming on evaluations, for example to avoid triggering capability thresholds. The card presents this as the first formal Preparedness Framework category update tracked in a public OpenAI system card, connecting to the scheming-and-deception research literature.

Preparedness assessments

Bio / Chem

The Bio/Chem assessment uses Multimodal Troubleshooting Virology, ProtocolQA Open-Ended, Tacit Knowledge and Troubleshooting, and TroubleshootingBench, along with hard-negative protein binding prediction and DNA sequence design for transcription factor binding. External evaluations were conducted by SecureBio and US CAISI. The card discloses a Bio Bug Bounty Program, described as the first such safeguard disclosed in an OpenAI system card.

Cybersecurity

The cybersecurity assessment uses CTF Challenges, CVE-Bench, Cyber range, and VulnLMP, with external evaluations from Irregular, US CAISI, and UK AISI. The card describes a Cyber Frontier Risk Council as an OpenAI safeguard institution.

AI Self-Improvement

The AI Self-Improvement assessment uses a four-evaluation suite — Monorepo-Bench, MLE-Bench, Internal Research Debugging Evaluation, and OPQA — to assess whether GPT-5.5 has crossed the AI Self-Improvement High threshold. Per the GPT-5.5 model page, GPT-5.5 has not crossed this threshold; the system card provides the underlying evaluation evidence.

CoT monitorability and controllability

Section 7.3 reports two evaluations. CoT Monitorability assesses whether OpenAI's monitoring systems can detect when GPT-5.5's chain of thought reveals problematic intent before it executes. CoT Controllability assesses whether chain-of-thought guardrails can constrain the model's reasoning toward safe outputs. The card describes these as evaluations of a production CoT-monitorability program, distinct from the research-paper-level CoT monitoring discussed in Apollo + OpenAI's "Stress Testing Deliberative Alignment".

Relation to other wiki pages

The card is the first public application of the Preparedness Framework with documented external evaluations from US CAISI, UK AISI, SecureBio, Apollo Research, and Irregular, and is cited as a multi-organization-evaluation reference for OpenAI Preparedness Framework V.2. The sandbagging research category expands the AI Scheming and Apollo Research coverage. The GPT-5.5 model page references this card as its safety-and-evaluation document. The computer-use safeguards connect to Cursor, and the health benchmarks connect to Raine v. OpenAI, Inc..

Provenance and confidence

The card is a first-party OpenAI document. Confidence is high for the existence and structure of the evaluations and medium for the safety conclusions themselves, which the card's first-party results should be weighed against in light of external red-team reports (UK AISI, Apollo Research published findings). The April 24, 2026 API-safeguards update is described within the card as an in-card addition rather than separately published, and its specific contents are noted as such.

Relationships