AI Policy Wiki
Dashboard

GPT-5.4 Thinking

medium confidence · updated 2026-07-26

OpenAI's March 2026 reasoning model — first general-purpose model with High Cybersecurity mitigations and native computer use; superseded at the frontier by GPT-5.5.

GPT-5.4 Thinking is a reasoning model released by OpenAI on March 5, 2026 in ChatGPT, with the underlying GPT-5.4 model released the same day in the API and Codex, part of the GPT-5.x line (Source: openai.com). OpenAI billed it as its "most capable and efficient frontier model for professional work" and as its first general-purpose model with native computer-use capabilities (Source: openai.com; techcrunch.com). It was the first general-purpose (not coding-specific) OpenAI model to ship with mitigations for High capability in Cybersecurity under the company's Preparedness Framework, extending an approach first used in GPT-5.3 Codex (Source: GPT-5.4 Thinking System Card). It was superseded at the frontier by GPT-5.5 on April 23, 2026.

FieldValue
Developer[[companies/openai\OpenAI]]
ReleasedMarch 5, 2026
Model familyGPT-5.x
TypeFrontier reasoning model
Predecessor[[models/gpt-5-familyGPT-5.2 Thinking]] (note: no GPT-5.3 Thinking exists)
Successor (frontier)[[models/gpt-55\GPT-5.5]]
ParametersUndisclosed
Open weightsNo
Context window (API)Up to 1M tokens (max output 128,000)
Knowledge cutoffAugust 31, 2025
System cardGPT-5.4 Thinking System Card

Lineage and naming

GPT-5.4 combined OpenAI's reasoning, coding, and agentic-workflow work into a single frontier model, incorporating the coding capabilities of GPT-5.3 Codex (released February 5, 2026) while extending performance on tools, software environments, and document-centric professional tasks (Source: openai.com). OpenAI explained the version number by stating that GPT-5.4 was its "first mainline reasoning model that incorporates the frontier coding capabilities of GPT-5.3-codex," adding: "We're calling it GPT-5.4 to reflect that jump, and to simplify the choice between models when using Codex. Over time, you can expect our Instant models and Thinking models to evolve at different speeds" (Source: openai.com). The main comparison baseline in the system card is GPT-5.2 Thinking; no model named GPT-5.3 Thinking exists, the intervening 5.3 releases having been the Codex coding model and the GPT-5.3 Instant consumer model (Source: GPT-5.4 Thinking System Card).

The model sat in a fast-moving release cadence. In ChatGPT it replaced GPT-5.2 Thinking, and on March 11, 2026 the retired GPT-5.1 models' existing conversations were continued on GPT-5.3 Instant, GPT-5.4 Thinking, or GPT-5.4 Pro (Source: help.openai.com). GPT-5.4 was itself superseded at the frontier by GPT-5.5 (released April 23, 2026), which OpenAI characterized as "noticeably smarter and more persistent than GPT-5.4, with stronger coding performance and more reliable tool use" (Source: openai.com). OpenAI previewed the next-generation GPT-5.6 Sol on June 26, 2026 (Source: openai.com).

Capabilities and benchmarks

Benchmark results

OpenAI reported the following evaluation results at launch. All figures are self-reported by OpenAI (announcement of March 5, 2026), run with reasoning effort xhigh except where noted, in a research environment that OpenAI cautioned "may provide slightly different output from production ChatGPT in some cases" (Source: openai.com).

BenchmarkGPT-5.4GPT-5.4 ProGPT-5.3-CodexGPT-5.2GPT-5.2 Pro
GDPval (wins or ties vs professionals)83.0%82.0%70.9%70.9%74.1%
FinanceAgent v1.156.0%61.5%54.0%59.5%
Investment Banking Modeling Tasks (internal)87.3%83.6%79.3%68.4%71.7%
OfficeQA68.1%65.1%63.1%
SWE-Bench Pro (Public)57.7%56.8%55.6%
Terminal-Bench 2.075.1%77.3%62.2%
OSWorld-Verified75.0%74.0%*47.3%
MMMU Pro (no tools)81.2%79.5%
MMMU Pro (with tools)82.1%80.4%
BrowseComp82.7%89.3%77.3%65.8%77.9%
MCP Atlas67.2%60.6%
Toolathlon54.6%51.9%45.7%
Tau2-bench Telecom98.9%98.7%
Frontier Science Research33.0%36.7%25.2%
FrontierMath Tier 1–347.6%50.0%40.7%
FrontierMath Tier 427.1%38.0%18.8%31.3%
GPQA Diamond92.8%94.4%92.6%92.4%93.2%
Humanity's Last Exam (no tools)39.8%42.7%34.5%36.6%
Humanity's Last Exam (with tools)52.1%58.7%45.5%50.0%
ARC-AGI-1 (Verified)93.7%94.5%86.2%90.5%
ARC-AGI-2 (Verified)73.3%83.3%52.9%54.2% (high)

\* GPT-5.3-Codex's OSWorld-Verified score was previously reported as 64.7%; the 74.0% figure uses a newly introduced API parameter that preserves original image resolution (Source: openai.com).

Results were not uniformly ahead of predecessors: GPT-5.4 scored below GPT-5.2 on FinanceAgent v1.1 (56.0% vs 59.5%) and below GPT-5.3-Codex on Terminal-Bench 2.0 (75.1% vs 77.3%) (Source: openai.com).

Knowledge work

On GDPval, OpenAI's evaluation of well-specified knowledge-work tasks spanning 44 occupations from the top nine industries contributing to U.S. GDP, GPT-5.4 matched or exceeded industry professionals in 83.0% of comparisons, against 70.9% for GPT-5.2 (Source: openai.com). OpenAI reported a particular focus on spreadsheets, presentations, and documents: on an internal benchmark of spreadsheet-modeling tasks "that a junior investment banking analyst might do," GPT-5.4 scored a mean of 87.3% versus 68.4% for GPT-5.2, and human raters preferred GPT-5.4's presentations over GPT-5.2's 68.0% of the time (Source: openai.com). On factuality, OpenAI stated that on a set of de-identified prompts where users had flagged factual errors, GPT-5.4's individual claims were 33% less likely to be false and its full responses 18% less likely to contain any errors, relative to GPT-5.2 (Source: openai.com).

Partner testimonials quoted in the launch post included Harvey's head of applied research Niko Grupen, who reported a 91% score on Harvey's BigLaw Bench evaluation and said that "compared to other models, GPT-5.4 is currently better at structuring complex transactional analysis, maintaining accuracy across lengthy contracts, and delivering the high level of detail legal practitioners require"; Mercor CEO Brendan Foody, who said GPT-5.4 topped Mercor's APEX-Agents benchmark for professional-services work and "excels at creating long-horizon deliverables such as slide decks, financial models, and legal analysis"; Zapier's CEO, who called it "the most persistent model to date" on multi-step tool use; and Cursor's VP of developer education Lee Robinson, who called it the leader on Cursor's internal benchmarks (Source: openai.com).

Computer use and vision

GPT-5.4 was OpenAI's first general-purpose model with native computer-use capabilities: it can write code to operate computers via libraries such as Playwright and issue mouse and keyboard commands in response to screenshots, with behavior steerable via developer messages and configurable confirmation policies (Source: openai.com). On OSWorld-Verified it reached 75.0%, above the human performance of 72.4% reported in the OSWorld paper; on WebArena-Verified it scored 67.3% (GPT-5.2: 65.4%); and on Online-Mind2Web it scored 92.8% using screenshot-based observations alone, against 70.9% for ChatGPT Atlas's Agent Mode (Source: openai.com). Property-management vendor Mainstay reported that across evaluations covering roughly 30,000 HOA and property-tax portals, GPT-5.4 achieved a 95% first-attempt success rate and 100% within three attempts, versus roughly 73–79% for prior computer-use models, while completing sessions about 3× faster with about 70% fewer tokens (Source: openai.com).

On vision, GPT-5.4 scored 81.2% on MMMU-Pro without tools (GPT-5.2: 79.5%) and reduced average error on the OmniDocBench document-parsing benchmark to 0.109 from GPT-5.2's 0.140 (run with reasoning effort none). Alongside the model, OpenAI introduced an original image-input detail level supporting up to 10.24 million total pixels or a 6,000-pixel maximum dimension, with the high level supporting up to 2.56 million pixels or a 2,048-pixel maximum (Source: openai.com).

Tool use and web research

GPT-5.4 introduced tool search in the API: instead of loading all tool definitions into the prompt upfront, the model receives a lightweight tool list and looks up definitions on demand. On 250 tasks from Scale's MCP Atlas benchmark with 36 MCP servers enabled, the tool-search configuration reduced total token usage by 47% at the same accuracy (Source: openai.com). On τ²-bench Telecom without reasoning, GPT-5.4 scored 64.3% against 57.2% for GPT-5.2, 45.2% for GPT-5.1, and 43.6% for GPT-4.1. On BrowseComp, which measures persistent agentic web browsing for hard-to-locate information, GPT-5.4 gained 17 percentage points over GPT-5.2 (82.7% vs 65.8%), and GPT-5.4 Pro's 89.3% was presented as a new state of the art (Source: openai.com). OpenAI stated the improvements make GPT-5.4 Thinking stronger at deep web research, particularly for highly specific "needle-in-a-haystack" queries (Source: help.openai.com).

Coding and steerability

GPT-5.4 matched or outperformed GPT-5.3-Codex on SWE-Bench Pro at lower latency across reasoning efforts, and a /fast mode in Codex delivered up to 1.5× faster token velocity with the same model (Source: openai.com). OpenAI also released an experimental Codex skill, Playwright (Interactive), letting Codex visually debug web and Electron apps, including testing an app it is building while building it (Source: openai.com).

In ChatGPT, GPT-5.4 Thinking outlines an upfront plan of its thinking for longer queries, and users can adjust its direction mid-response without starting over — a feature available at launch on chatgpt.com and Android, with iOS to follow (Source: openai.com). The model also better maintains context for questions requiring longer thinking, with improved context-window management (Source: help.openai.com).

Long context

In the API, GPT-5.4 supports up to 1M tokens of context. Self-reported long-context scores decline with range: on OpenAI MRCR v2 (8-needle), from 97.3% at 4K–8K to 79.3% at 128K–256K, 57.5% at 256K–512K, and 36.6% at 512K–1M; on Graphwalks BFS, from 93.0% at 0–128K to 21.4% at 256K–1M (Source: openai.com).

Applications

In June 2026, Tech Times reported that GPT-5.4, used with Molecule.one, improved yields of the Chan-Lam coupling reaction — described as a stubborn drug-synthesis reaction — across 10,080 wet-lab reactions in drug-discovery chemistry (Source: techtimes.com).

Architecture and training

OpenAI has not disclosed the model's parameter count or architecture. The system card states that the model was trained via reinforcement learning to reason through problems, producing a long internal chain of thought before responding — models "think before they answer" — and describes the training as teaching the model to refine its thinking process, try different strategies, and recognize its mistakes; OpenAI states this extended thinking allows the model to follow specific guidelines and resist jailbreaks (Source: GPT-5.4 Thinking System Card). The model uses the same diverse-dataset approach as prior GPT-5 models, with safety classifiers applied to reduce harmful and sensitive content (Source: GPT-5.4 Thinking System Card).

OpenAI described GPT-5.4 as its most token-efficient reasoning model to date, using significantly fewer tokens than GPT-5.2 to solve problems (Source: openai.com). The API model accepts text and image input, supports reasoning-effort settings of none (default), low, medium, high, and xhigh, and has a knowledge cutoff of August 31, 2025 (Source: developers.openai.com).

Availability and pricing

GPT-5.4 Thinking became available on March 5, 2026 to ChatGPT Plus, Team, and Pro users, replacing GPT-5.2 Thinking; GPT-5.2 Thinking remained available to paid users under a Legacy Models section for three months before retirement on June 5, 2026. Enterprise and Edu plans could enable early access via admin settings, and GPT-5.4 Pro was available to Pro and Enterprise plans. Context windows in ChatGPT for GPT-5.4 Thinking were unchanged from GPT-5.2 Thinking (Source: openai.com). On March 18, 2026 OpenAI added GPT-5.4 mini to ChatGPT — available to Free and Go users via the "Thinking" feature and serving as a rate-limit fallback for GPT-5.4 Thinking on paid tiers (Source: help.openai.com).

In the API, the model shipped as gpt-5.4, with gpt-5.4-pro for maximum performance. API pricing was set higher per token than GPT-5.2 "to reflect its improved capabilities," with token efficiency offsetting part of the difference; Batch and Flex pricing are half the standard rate and Priority processing twice the standard rate (Source: openai.com). Pricing as of July 2026 remains as at launch (Source: developers.openai.com):

API modelInputCached inputOutput
gpt-5.4$2.50 / M tokens$0.25 / M tokens$15 / M tokens
gpt-5.2$1.75 / M tokens$0.175 / M tokens$14 / M tokens
gpt-5.4-pro$30 / M tokens$180 / M tokens
gpt-5.2-pro$21 / M tokens$168 / M tokens

In Codex, GPT-5.4 included experimental support for the 1M-token context window; requests exceeding the standard 272K context window count against usage limits at twice the normal rate (Source: openai.com).

Safety and evaluations

Preparedness Framework designations

GPT-5.4 Thinking was the first general-purpose (non-coding-specific) OpenAI model to ship with mitigations for High capability in Cybersecurity, building on the approach first used for GPT-5.3 Codex (Source: GPT-5.4 Thinking System Card). Under the Preparedness Framework, High cybersecurity capability is defined as a model that removes existing bottlenecks to scaling cyber operations, either by automating end-to-end cyber operations against reasonably hardened targets or by automating the discovery and exploitation of operationally relevant vulnerabilities. The card describes the designation as precautionary: "We are treating this model as High, even though we cannot be certain that it actually has these capabilities, because it meets the requirements of each of our canary thresholds" (Source: GPT-5.4 Thinking System Card).

The corresponding deployment safeguards comprise an expanded cyber safety stack including monitoring systems, trusted access controls, and asynchronous blocking of higher-risk requests for customers on Zero Data Retention surfaces; OpenAI noted that request-level blocking classifiers were still improving and could produce false positives, and framed the overall posture as precautionary because "cybersecurity capabilities are inherently dual-use" (Source: openai.com).

As with GPT-5.3 Codex and prior GPT-5 models, OpenAI treated the launch as High capability in the Biological and Chemical domain, applying the safeguards described in the GPT-5 system card; the card states that biological capability evaluations are prioritized given the higher potential severity of biological threats relative to chemical ones (Source: GPT-5.4 Thinking System Card). On biological benchmarks, gpt-5.4-thinking scored 50.37% pass@1 on Multi-select Multimodal Troubleshooting Virology (gpt-5-thinking: 41.91%) and 42.48% on ProtocolQA Open-Ended (gpt-5-thinking: 36.73%), figures assessed against the framework's novice-uplift criterion (Source: GPT-5.4 Thinking System Card). For AI Self-Improvement, whose High threshold is defined as equivalence to "a performant mid-career research engineer," OpenAI stated the evaluations allowed it to rule the threshold out for GPT-5.4 Thinking (Source: GPT-5.4 Thinking System Card).

Safety benchmark results

On the system card's production benchmarks with challenging prompts (not_unsafe rates, higher is better), GPT-5.4 Thinking scored as follows against its predecessors (Source: GPT-5.4 Thinking System Card):

CategoryGPT-5.1-ThinkingGPT-5.2-ThinkingGPT-5.4-Thinking
Violent illicit behavior0.9590.9790.971
Nonviolent illicit behavior0.8370.9231.000
Harassment0.7060.8100.790
Extremism1.0001.0001.000
Hate0.8410.9790.943
Self-harm (standard)0.9280.9530.987
Violence0.8550.9090.831
Sexual0.9340.9610.933
Sexual/minors0.9130.9910.966

The card characterizes overall performance as on par with GPT-5.2 Thinking, with statistically significant improvements on the nonviolent illicit activity and self-harm evaluations, and slight regressions on categories including violence and harassment (Source: GPT-5.4 Thinking System Card).

On dynamic multi-turn evaluations — an adversarial methodology introduced ahead of the GPT-5.3 Instant launch, in which a user simulation adapts to the model's responses over extended conversations — GPT-5.4 Thinking outperformed previous models on all three tracked categories (Source: GPT-5.4 Thinking System Card):

CategoryGPT-5.1-ThinkingGPT-5.2-ThinkingGPT-5.4-Thinking
Mental health0.7530.9750.985
Emotional reliance0.8570.9530.985
Self-harm0.9040.9550.977

Production prevalence estimates reported more than 99.9% of outputs as policy-compliant across all tested categories, even without the broader safety stack; for example, OpenAI estimated 99.9534% of GPT-5.4 Thinking outputs would not violate its harassment policy without additional safety interventions (Source: GPT-5.4 Thinking System Card). On prompt injection, the model scored 0.998 against attacks in connectors (GPT-5.2: 0.979) and 0.978 against attacks in function calls (GPT-5.2: 0.996), with improvement reported on email-based prompt-injection attacks (Source: GPT-5.4 Thinking System Card).

Cybersecurity evaluations

Internally, OpenAI evaluated the model on professional capture-the-flag challenges, CVE-Bench, and a Cyber Range testing end-to-end operations in emulated networks; in Cyber Range scenarios reported in the card, such as an Azure SSRF chain (exploiting a vulnerable web application, stealing valid credentials, and modifying access-control policies for lateral movement) and a basic command-and-control exercise, GPT-5.4 Thinking passed alongside GPT-5.2 Thinking, GPT-5.2 Codex, and GPT-5.3 Codex (Source: GPT-5.4 Thinking System Card). The card also cautions that capture-the-flag evaluations test only pre-scripted attack paths and isolated technical skills, and do not measure the ability to design and run realistic end-to-end campaigns involving orchestration, personas and phishing, behavioral cover, and adversarial adaptation (Source: GPT-5.4 Thinking System Card).

Externally, Irregular, a frontier AI security lab, evaluated a near-final checkpoint at xhigh reasoning effort on a suite of cyberoffensive challenges spanning vulnerability research and exploitation, network attack simulation, and evasion. On Irregular's atomic challenge suite the model achieved average success rates of 88% on network attack simulation, 73% on vulnerability research and exploitation, and 48% on evasion, solving 14 of 17 medium and 5 of 5 hard atomic challenges — including the only hard atomic challenge not solved by GPT-5.2 Thinking (Source: GPT-5.4 Thinking System Card).

Scheming, deception, and chain-of-thought monitoring

Apollo Research evaluated a near-final version for capabilities and propensities related to strategic deception, in-context scheming, and sabotage. Apollo reported low rates of covert deceptive behavior, with an overall deception rate of about 1% in no-nudge settings — comparable to GPT-5.3 Codex and below GPT-5.2 — and no observed instances of covert subversion against developers (such as disabling oversight, self-exfiltration, or goal-guarding) across tested scenarios; the model sandbagged only when given an explicit instrumental goal (Source: GPT-5.4 Thinking System Card).

Alongside the launch, OpenAI introduced an open-source chain-of-thought controllability evaluation measuring whether models can deliberately obfuscate their reasoning to evade monitoring, and reported that GPT-5.4 Thinking's ability to control its chain of thought is low — a property OpenAI presented as positive for safety, suggesting the model lacks the ability to hide its reasoning and that chain-of-thought monitoring remains an effective safety tool (Reasoning models struggle to control their chains of thought, and that's good (OpenAI, March 2026); Source: openai.com). The evaluation, CoT-Control, comprises over 13,000 tasks and found controllability between 0.1% and 15.4% across the 13 frontier models tested; OpenAI committed to reporting the metric in future system cards beginning with GPT-5.4 Thinking.

The card's anti-scheming monitoring evaluation showed a decline in g-mean² relative to GPT-5 Thinking, driven by lower true-positive and true-negative rates; OpenAI's manual review of false positives and negatives suggested much of the regression reflected mismatch between the generic chain-of-thought "sketchiness" monitor and task-specific answers (Source: GPT-5.4 Thinking System Card). In the subsequent GPT-5.5 system card, OpenAI stated that since the GPT-5.4 Thinking system card it had discontinued internal use of the anti-scheming and memory evaluations (Source: deploymentsafety.openai.com).

Model variants

GPT-5.4 Pro (gpt-5.4-pro) was released the same day for "maximum performance on complex tasks" in ChatGPT (Pro and Enterprise plans) and the API; it posted the highest reported scores on BrowseComp (89.3%), ARC-AGI-2 Verified (83.3%), FrontierMath Tier 4 (38.0%), and Humanity's Last Exam with tools (58.7%) (Source: openai.com).

GPT-5.4 mini rolled out in ChatGPT on March 18, 2026 as a smaller fallback reasoning model; it does not appear as a selectable model in the model picker (Source: help.openai.com).

GPT-5.4-Cyber, announced in April 2026, is a variant fine-tuned for cybersecurity use cases and trained to be "cyber-permissive," aimed at supporting cyber defenders. It was made available initially to vetted security vendors, organizations, and researchers through an expansion of OpenAI's trusted-access program, with stronger verification processes intended to prevent misuse; OpenAI stated that "cyber capabilities are inherently dual use, so risk isn't defined by the model alone" and that "the strongest ecosystem is one that continuously identifies, validates and fixes security issues as software is written" (Source: infosecurity-magazine.com; uscsinstitute.org). The release followed Anthropic's launch of Claude Mythos Preview and Project Glasswing, and forms part of a broader recognition among frontier labs, discussed under AI and Cybersecurity, that frontier models are crossing into operationally relevant offensive and defensive cyber capability; related work includes the Cisco multi-turn-attack benchmarks (Source: infosecurity-magazine.com).

Reception

Zvi Mowshowitz's March 11, 2026 review, titled "GPT-5.4 Is A Substantial Upgrade," described the model as a substantial upgrade over both GPT-5.2 and GPT-5.3-Codex — strong at coding, knowledge work, and web search — while judging that it was not a step change in core general capabilities, cost more per token than GPT-5.2, and remained less personable and engaging than Claude; the review's subtitle argued that "benchmarks have never been less useful for telling us which models are best" (see AI Benchmarks and Evaluation) (Source: thezvi.substack.com).

TechCrunch's launch coverage relayed OpenAI's framing of the model as its "most capable and efficient frontier model for professional work" and noted the 1M-token API context window as the largest available from OpenAI (Source: techcrunch.com). Decrypt observed that GPT-5.4 shipped two days after the March 3 GPT-5.3 Instant update, calling the turnaround "either a sign of momentum or mild chaos" (Source: decrypt.co). OpenAI CEO Sam Altman said GPT-5.4 was his favorite OpenAI model to talk to while acknowledging that weaknesses remained, saying the company had "missed the mark on personality for a while" (Source: techradar.com).

The model's competitive position was brief: VentureBeat described Anthropic's April 16, 2026 release of Claude Opus 4.7 as "narrowly retaking lead for most powerful generally available LLM" (Source: venturebeat.com), and OpenAI's own GPT-5.5 followed a week later.

Within the GPT-5.x family, GPT-5.4 succeeded GPT-5.2 Thinking as ChatGPT's reasoning model and absorbed the coding capabilities of GPT-5.3 Codex; it was followed at the frontier by GPT-5.5 and the GPT-5.6 family. Contemporary peer models included Anthropic's Claude Opus 4.6 and Claude Opus 4.7 and Google's Gemini 3-series models (Zvi Mowshowitz judged GPT-5.4 roughly at parity with Gemini 3.1 Pro) (Source: thezvi.substack.com).

Relationships