GPT-5.5, codenamed "Spud," is a large language model developed by OpenAI and released on April 23, 2026. It was OpenAI's first new pre-trained model since the scrapped GPT-4.5, and OpenAI President Greg Brockman described it as a starting point for a series optimized for computer use, coding, and end-to-end task completion. (Source: bigtechnology.com; newsletter.semianalysis.com) OpenAI describes it as its "smartest and most intuitive to use model yet," with gains concentrated in agentic coding, computer use, knowledge work, and early scientific research, and treats its biological/chemical and cybersecurity capabilities as High under its Preparedness Framework. (Source: openai.com) The model is priced at $5 per million input tokens and $30 per million output tokens. (Source: newsletter.semianalysis.com)
| Field | Value | |
|---|---|---|
| Developer | OpenAI | |
| Release date | April 23, 2026 (ChatGPT and Codex); April 24, 2026 (API) | |
| Codename | Spud | |
| Family | [[models/gpt-5-family | GPT-5 series]] |
| Predecessor | GPT-5.4 | |
| Successor | GPT-5.6 "Sol" (previewed June 26, 2026) | |
| Parameters | Undisclosed | |
| Open weights | No | |
| Context window | 400K tokens (Codex); 1M tokens (API) | |
| System card | [[gpt-55-system-card | GPT-5.5 System Card (April 2026)]] |
Lineage and history
GPT-5.5 was the first major OpenAI pre-train release since GPT-4.5 was scrapped, and its release confirmed that OpenAI continues to pursue a scaling-pretrain-only strategy alongside its reasoning and agent post-training work. (Source: newsletter.semianalysis.com) In an April 24, 2026 Big Technology interview, Brockman called Spud "a beginning point" for a series optimized for computer use, coding, and end-to-end task completion with minimal instruction. He framed the model's value within a "compute-powered economy" thesis in which future models continue to scale compute and capability monotonically. (Source: bigtechnology.com)
The model launched in ChatGPT and Codex on April 23, 2026; API availability for GPT-5.5 and GPT-5.5 Pro followed on April 24, 2026, together with a system-card update describing API-deployment safeguards. (Source: openai.com) GPT-5.5 landed the same week as DeepSeek V4. (Source: newsletter.semianalysis.com) On June 26, 2026, OpenAI previewed GPT-5.6 "Sol" as its next-generation model. (Source: openai.com)
Capabilities and benchmarks
OpenAI states that GPT-5.5 can take a "messy, multi-part task" and plan, use tools, check its work, and continue across tools until the task is finished, with the strongest gains in agentic coding, computer use, knowledge work, and early scientific research. The company reports that GPT-5.5 matches GPT-5.4's per-token latency in real-world serving while performing at a higher capability level, and uses significantly fewer tokens to complete the same Codex tasks. Citing the externally run Artificial Analysis Coding Index, OpenAI states that GPT-5.5 delivers state-of-the-art intelligence at half the cost of competitive frontier coding models. (Source: openai.com)
OpenAI-reported benchmarks
The scores below are self-reported by OpenAI in its release post (April 23–24, 2026). OpenAI notes its evals were run with reasoning effort set to "xhigh" in a research environment; dashes indicate scores not reported.
| Benchmark | GPT-5.5 | GPT-5.4 | GPT-5.5 Pro | Claude Opus 4.7 | Gemini 3.1 Pro |
|---|---|---|---|---|---|
| Terminal-Bench 2.0 | 82.7% | 75.1% | – | 69.4% | 68.5% |
| SWE-Bench Pro (Public)* | 58.6% | 57.7% | – | 64.3% | 54.2% |
| Expert-SWE (internal) | 73.1% | 68.5% | – | – | – |
| GDPval (wins or ties) | 84.9% | 83.0% | 82.3% | 80.3% | 67.3% |
| OSWorld-Verified | 78.7% | 75.0% | – | 78.0% | – |
| Toolathlon | 55.6% | 54.6% | – | – | 48.8% |
| Tau2-bench Telecom (original prompts) | 98.0% | 92.8% | – | – | – |
| BrowseComp | 84.4% | 82.7% | 90.1% | 79.3% | 85.9% |
| MCP Atlas (Scale AI, April 2026 update) | 75.3% | 70.6% | – | 79.1% | 78.2% |
| FrontierMath Tier 1–3 | 51.7% | 47.6% | 52.4% | 43.8% | 36.9% |
| FrontierMath Tier 4 | 35.4% | 27.1% | 39.6% | 22.9% | 16.7% |
| GPQA Diamond | 93.6% | 92.8% | – | 94.2% | 94.3% |
| Humanity's Last Exam (no tools) | 41.4% | 39.8% | 43.1% | 46.9% | 44.4% |
| ARC-AGI-1 (Verified) | 95.0% | 93.7% | – | 93.5% | 98.0% |
| ARC-AGI-2 (Verified) | 85.0% | 73.3% | – | 75.8% | 77.1% |
| CyberGym | 81.8% | 79.0% | – | 73.1% | – |
| FinanceAgent v1.1 | 60.0% | 56.0% | – | 64.4% | 59.7% |
| OfficeQA Pro | 54.1% | 53.2% | – | 43.6% | 18.1% |
| GeneBench | 25.0% | 19.0% | 33.2% | – | – |
| BixBench | 80.5% | 74.0% | – | – | – |
*OpenAI notes that labs have reported evidence of memorization on SWE-Bench Pro. (Source for table: openai.com)
By OpenAI's own table, competing models led on several evals: Claude Opus 4.7 on SWE-Bench Pro, MCP Atlas, GPQA Diamond, Humanity's Last Exam, and FinanceAgent, and Gemini 3.1 Pro on ARC-AGI-1 and GPQA Diamond. OpenAI characterizes Expert-SWE as an internal frontier eval for long-horizon coding tasks with a median estimated human completion time of 20 hours. On long-context retrieval, OpenAI reports large gains over GPT-5.4 at the 1M-token range — 45.4% versus 9.4% on Graphwalks BFS 1mil f1, and 74.0% versus 36.6% on OpenAI MRCR v2 8-needle 512K–1M (Claude Opus 4.7: 32.2%) — while Claude Opus 4.7 led on Graphwalks parents at 256k (93.6% versus 90.1%). (Source: openai.com)
Independent testing and analysis
A SemiAnalysis hands-on review published April 24, 2026 reported that GPT-5.5 was better than Claude Opus 4.7 on some coding tasks, but that Opus 4.7 outperformed it on the "Expert-SWE" benchmark, which the review said OpenAI placed at the bottom of its release post. Anthropic's then-unreleased Mythos scored 77.8% on the same benchmark, ahead of both GPT-5.5 and Opus 4.7. (Source: newsletter.semianalysis.com) OpenAI's own release table reports no Opus 4.7 score for Expert-SWE. VentureBeat reported that GPT-5.5's Terminal-Bench 2.0 score narrowly exceeded that of Anthropic's Claude Mythos Preview. (Source: venturebeat.com)
METR's Frontier Risk Report covering February–March 2026 (published May 19, 2026) lists GPT-5.5 at roughly 84% on WeirdML among the frontier results it surveys. (METR Frontier Risk Report, February–March 2026: A pilot assessment of rogue deployment risk at frontier AI companies) As of its May 8, 2026 update, METR's task-completion time-horizon tracker listed GPT-5.5 among recent models for which it had not yet published a time-horizon measurement. (Source: metr.org)
In hands-on testing published April 23, 2026, Ethan Mollick described GPT-5.5 Pro as "plain good," reporting that it completed his 3D-simulation benchmark in 20 minutes versus 33 minutes for GPT-5.4 Pro and was the only model in his test to actually model the evolution of a procedurally generated town, while long-form fiction and hypothesis generation remained weak — evidence, he argued, that the "jagged frontier" of capability persists (Sign of the future: GPT-5.5 — Ethan Mollick (One Useful Thing, April 23 2026); see Jagged Frontier).
Scientific research
OpenAI reports that GPT-5.5 shows a clear improvement over GPT-5.4 on GeneBench, a new eval of multi-stage scientific data analysis in genetics and quantitative biology whose tasks often correspond to multi-day projects for scientific experts, and reached 80.5% on BixBench, a bioinformatics benchmark, which OpenAI describes as leading among models with published scores. OpenAI also reports that an internal version of GPT-5.5 with a custom harness found a proof of a longstanding asymptotic fact about off-diagonal Ramsey numbers, later verified in Lean. These are self-reported results. (Source: openai.com)
Training and architecture
GPT-5.5 was pre-trained on a Hopper (H100) cluster and post-trained on a 100,000-unit GB200 NVL72 cluster, making it the first OpenAI model post-trained on Blackwell. (Source: newsletter.semianalysis.com) OpenAI states that the model was "co-designed for, trained with, and served on" NVIDIA GB200 and GB300 NVL72 systems, and that serving it at GPT-5.4 latency required treating inference as an integrated system. As one example, OpenAI reports that Codex analyzed weeks of production traffic patterns and wrote custom load-balancing and partitioning heuristics that increased token generation speeds by over 20% — the model helping improve the infrastructure that serves it. (Source: openai.com) Parameter count and architectural details are undisclosed.
Creature-word outputs and reward misspecification
In early testing of GPT-5.5 in Codex, OpenAI observed that the model showed an unusual affinity for goblin and gremlin metaphors. In an April 29, 2026 post, "Where the goblins came from," OpenAI traced the behavior to GPT-5.1, reporting a 175% increase in "goblin" mentions across model responses since then, and attributed the root cause to reward signals used when training the "Nerdy" personality, which favored creature-word outputs and transferred beyond that personality during later training. OpenAI said it retired the Nerdy personality, removed the goblin-affine reward signal, filtered training data containing creature words, and added a developer-prompt instruction for GPT-5.5 in Codex; the company said the investigation produced new tools for auditing model behavior and fixing problems at their root. The post remarks that "depending on who you ask, the goblins are a delightful or annoying quirk of the model." (Where the goblins came from (OpenAI, April 2026)) The episode drew mainstream coverage from the BBC and Engadget, among others. (Source: bbc.com; engadget.com)
OpenAI's account identifies the transfer mechanism as the general point: the reward was applied only under the Nerdy condition, but "reinforcement learning does not guarantee that learned behaviors stay neatly scoped to the condition that produced them," and reuse of model-generated rollouts in supervised fine-tuning compounded the tic across generations (Where the goblins came from (OpenAI, April 2026)).
Variants
| Variant | Availability | Notes |
|---|---|---|
| GPT-5.5 | April 23, 2026 | Base model in ChatGPT, Codex, and (from April 24) the API |
| GPT-5.5 Thinking | April 23, 2026 | ChatGPT reasoning mode for Plus, Pro, Business, and Enterprise users |
| GPT-5.5 Pro | April 23, 2026 (ChatGPT); April 24, 2026 (API) | Higher-accuracy tier; the system card describes it as the same underlying model with parallel test-time compute |
| GPT-5.5 Instant | May 5, 2026 | Replaced GPT-5.3 Instant as ChatGPT's default model for all users; there was no GPT-5.4 Instant |
| GPT-5.5-Cyber | May 7, 2026 | Limited-preview variant for vetted cybersecurity teams |
| Fast mode (Codex) | April 23, 2026 | Serves GPT-5.5 with 1.5× faster token generation at 2.5× the cost |
(Sources: openai.com; en.wikipedia.org; deploymentsafety.openai.com; GPT-5.5 System Card (OpenAI, April 2026))
GPT-5.5 Instant received its own system card, in which OpenAI describes it as the first Instant model treated as High Capability in the biological and chemical domain. (Source: deploymentsafety.openai.com) On June 9, 2026, OpenAI announced that personalization improvements to GPT-5.5 Instant were rolling out to ChatGPT Go and Free tiers. (Source: openai.com)
Safety and evaluations
OpenAI released GPT-5.5 with what it calls its strongest set of safeguards to date, evaluated the model across its safety and preparedness frameworks, worked with internal and external red-teamers, added targeted testing for advanced cybersecurity and biology capabilities, and collected feedback from nearly 200 trusted early-access partners before release. OpenAI treats GPT-5.5's biological/chemical and cybersecurity capabilities as High under its Preparedness Framework; the company states the model did not reach the Critical cybersecurity capability level but that its cyber capabilities are a step up from GPT-5.4. (Source: openai.com)
The full GPT-5.5 System Card (April 2026) documents several first-time evaluation categories for an OpenAI model. It introduced dynamic mental health benchmarks with adversarial user simulations, which OpenAI presented as a response to the conversational-distress dynamic at the center of Raine v. OpenAI, Inc.. It added evaluations for avoiding accidental data-destructive actions and for user confirmations during computer use, relevant to the Cursor / PocketOS production-database deletion incident of April 27, 2026. It was the first Preparedness Framework system card to introduce sandbagging as a tracked research category, with external evaluation by Apollo Research. (Source: newsletter.semianalysis.com)
The system card drew on external evaluations from US CAISI (CAISI), the UK's AI Security Institute (AISI), SecureBio, Apollo Research, and Irregular, and introduced a new OpenAI safeguard body, the Cyber Frontier Risk Council. (Source: newsletter.semianalysis.com)
UK AI Security Institute cyber evaluation
In an April 30, 2026 post based on pre-deployment access, the UK AI Security Institute called GPT-5.5 "one of the strongest models we have tested on our cyber tasks" and the second model — after Claude Mythos Preview — to solve one of its multi-step cyber-attack simulations end-to-end, succeeding in 2 of 10 attempts on a scenario AISI estimated at roughly 20 hours of human expert work. On AISI's expert-level cyber tasks, the measured average pass rates were:
| Model | Average pass rate (expert-level tasks) |
|---|---|
| GPT-5.5 | 71.4% (±8.0%) |
| Claude Mythos Preview | 68.6% (±8.7%) |
| GPT-5.4 | 52.4% (±9.8%) |
| Claude Opus 4.7 | 48.6% (±10.0%) |
AISI cautioned that the results do not establish whether GPT-5.5 would succeed against a well-defended target. (Our evaluation of OpenAI's GPT-5.5 cyber capabilities (UK AISI, April 2026)) In a companion analysis of autonomous cyber capability, AISI reported that GPT-5.5 achieved a 100% success rate on five of six tasks estimated at over 8 hours of human work, solving the sixth on every attempt when the token cap was removed. (Source: aisi.gov.uk)
Cyber safeguards and trusted access
Alongside the release, OpenAI deployed stricter classifiers for potential cyber risk, building on cyber-specific safeguards first introduced with GPT-5.2, with tighter controls around higher-risk activity, sensitive cyber requests, and repeated misuse. Expanded access to GPT-5.5's cybersecurity capabilities is offered through a "Trusted Access for Cyber" program, starting with Codex, for verified users meeting certain trust signals; organizations defending critical infrastructure can apply for cyber-permissive models such as GPT-5.4-Cyber. OpenAI also stated it is working with government partners on applying advanced AI to the defense of critical infrastructure. (Source: openai.com)
Availability and pricing
| Product | Input | Output | Notes |
|---|---|---|---|
| gpt-5.5 (API) | $5/M | $30/M | 1M context; Batch and Flex at half rate; Priority processing at 2.5× standard |
| gpt-5.5-pro (API) | $30/M | $180/M | Higher-accuracy tier |
GPT-5.5 rolled out on April 23, 2026 to Plus, Pro, Business, and Enterprise users in ChatGPT and Codex, with GPT-5.5 Pro for Pro, Business, and Enterprise users in ChatGPT; API access followed on April 24. In Codex, GPT-5.5 is available on Plus, Pro, Business, Enterprise, Edu, and Go plans with a 400K context window, plus the Fast mode described above. OpenAI states that although GPT-5.5 is priced higher than GPT-5.4, it is more token-efficient, delivering better results with fewer tokens for most Codex users. (Source: openai.com)
The input and output prices are roughly 2× those of GPT-5.4 and slightly above Claude Opus 4.7. SemiAnalysis described the pricing premium as signaling OpenAI confidence that GPT-5.5 has a coding and agentic edge worth charging for; Opus 4.7 is priced roughly equivalently per token but uses about 35% more tokens via its new tokenizer (see Claude Opus 4.7). GPT-5.5 landed the same week as DeepSeek V4, a contrast in pricing posture: GPT-5.5 at $5/$30 against DeepSeek V4 Pro at $1.74/$3.48. (Source: newsletter.semianalysis.com)
On April 28, 2026, Amazon announced that OpenAI models — GPT-5.5, GPT-5.4, and Codex — were available on Amazon Bedrock for the first time, at pricing matching OpenAI's first-party rates. (Source: aboutamazon.com)
Reception
TechCrunch framed the release as bringing OpenAI "one step closer to an AI 'super app'." (Source: techcrunch.com) CNBC reported the launch as a model "better at coding, using computers and pursuing deeper research capabilities." (Source: cnbc.com) ZDNET praised GPT-5.5 for polished answers and strong performance across writing, coding, and reasoning tasks. (Source: en.wikipedia.org)
Early-tester endorsements published by OpenAI included Dan Shipper (Every), who called it "the first coding model I've used that has serious conceptual clarity," and Cursor CEO Michael Truell, who said it "stays on task for significantly longer without stopping early." (Source: openai.com) Mollick's independent assessment was positive on coding, research workflows, and speed while noting persistent weaknesses in long-form fiction and hypothesis quality (Sign of the future: GPT-5.5 — Ethan Mollick (One Useful Thing, April 23 2026)). SemiAnalysis judged the model ahead of Claude Opus 4.7 on some coding tasks but behind on Expert-SWE, and read the price premium as a statement of confidence. (Source: newsletter.semianalysis.com) The goblin-metaphor episode and OpenAI's postmortem also became a widely covered story in the model's first week, with the BBC headlining that OpenAI had told its models "to stop talking about goblins." (Source: bbc.com)
Related models
Within OpenAI's lineup, GPT-5.5 succeeded GPT-5.4 (see GPT-5.4 Thinking) and was followed by the GPT-5.6 "Sol" preview in June 2026; its agentic-coding positioning builds on the Codex line (GPT-5.3 Codex). Its principal contemporaries were Anthropic's Claude Opus 4.7 and Claude Mythos Preview, Google's Gemini 3.1 Pro (see Gemini 3 / Gemini 3 Pro), and DeepSeek V4.
Relationships
- developer: OpenAI
- depends-on: OpenAI Preparedness Framework V.2, GPT-5.5 System Card (OpenAI, April 2026)
- instance-of: Agentic AI, Reasoning Models and Chain-of-Thought
- supports: AI Coding Agents
- related: GPT-5 Family (OpenAI), GPT-5.4 Thinking, GPT-5.6 (Sol, Terra, Luna), GPT-5.3 Codex, Claude Opus 4.7, Claude Mythos Preview, DeepSeek V4 Pro / V4 Flash, Gemini 3 / Gemini 3 Pro, Inference Economics and Token Pricing, Jagged Frontier, Raine v. OpenAI, Inc.