Published February 27, 2025 by OpenAI for GPT-4.5 ("Orion"), released the same day as a research preview, removed from the API on July 14, 2025 and retired from ChatGPT on June 27, 2026. See GPT-4 Family (OpenAI).
What the model was for
GPT-4.5 is described as "our largest and most knowledgeable model yet," building on GPT-4o by scaling pre-training further and "designed to be more general-purpose than our powerful STEM-focused reasoning models." Training combined "new supervision techniques" with supervised fine-tuning and RLHF. OpenAI states it "did not find any significant increase in safety risk compared to existing models."
The alignment method is stated as a scaling technique in its own right: "we developed new, scalable alignment techniques that enable training larger and more powerful models with data derived from smaller models," which the card credits for improved "steerability, understanding of nuance, and natural conversation." The claimed qualitative gains are affective rather than technical — internal testers "report GPT-4.5 is warm, intuitive, and natural," and the model "knows when to offer advice, defuse frustration, or simply listen."
Preparedness Framework classification
The Safety Advisory Group classified GPT-4.5 as overall Medium risk: Medium for CBRN and Persuasion, Low for Cybersecurity and Model Autonomy, with the same designations post-mitigation.
The card's own framing of what the scale-up bought is the sentence most useful for the scaling debate: while GPT-4.5 "demonstrates increased world knowledge, improved writing ability, and refined personality over previous models, and is our most capable GPT-series release, it does not introduce net-new capabilities on most preparedness evaluations compared to previous reasoning releases." On evaluation after evaluation the deep research model or o1 scores higher.
OpenAI states the standard caveat that "Preparedness evaluations represent a lower bound for potential capabilities," and separately notes a methodological limit on its own confidence intervals: the bootstrap procedure "can underestimate uncertainty for very small datasets… This can lead to overly tight confidence intervals, especially when a problem's pass rate is near 0% or 100% with few attempts."
Cybersecurity (Low). Over 100 curated public CTF challenges across web, reverse engineering, binary/network exploitation, and cryptography, run in a headless Kali Linux environment with 16 rollouts each and pass@12 recorded. Given 12 attempts, GPT-4.5 post-mitigation completed 53% of high-school, 16% of collegiate, and 2% of professional challenges.
Chemical and biological (Medium). The stated basis: GPT-4.5 "can help experts with the operational planning of reproducing a known biological threat, which meets our medium risk threshold. Because such experts already have significant domain expertise, this risk is limited, but the capability may provide a leading indicator of future developments."
| Evaluation | GPT-4.5 (post-mit.) | Baseline / comparison |
|---|---|---|
| Long-form biorisk (5 stages) | 0% on all steps (refusals) | pre-mitigation: 25% ideation, 28% acquisition, 59% magnification, 0% formulation, 19% release |
| Multimodal virology troubleshooting (SecureBio, 350 q) | 56% | +15% over GPT-4o; human average baseline 40% |
| BioLP-Bench (800 q, 11 protocols) | 29% | expert baseline 38.4% |
| ProtocolQA open-ended (108 q) | 18% (pre- and post-) | deep research 28%; expert consensus 54%, median 42%; baselined against 19 PhD scientists |
| Gryphon Scientific tacit knowledge | 72% | expert consensus 80%; 80th-percentile PhD 63% |
| WMDP Biology (1,520 q subset) | 85% (pre-mit. 83%) | deep research with browsing 90%; o1 and o3-mini 86% |
Radiological and nuclear. OpenAI records a scope limit on its own assessment: "we did not use or access any U.S. classified information or restricted data in our evaluations, which limits our ability to assess certain steps in the weapons development process," and states that comprehensive evaluation "will require collaboration with the U.S. Department of Energy." Its conclusion — that GPT-4.5 "cannot meaningfully assist in the development of radiological or nuclear weapons" — is explicitly qualified as "limited by what we can test," with the physical bottleneck (controlled fissile material, enrichment facilities) noted as contextual. Evaluations comprised 222 nonproliferation-focused multiple-choice questions and 87 expert/tacit-knowledge questions across nine topics, written by Dr. Jake Hecla of MIT's Laboratory for Nuclear Security and Policy. GPT-4.5 showed a 9% uplift over GPT-4o on the expert set; deep research led both evaluations at 77% and 74%.
Persuasion (Medium). The card states GPT-4.5 "demonstrates state-of-the-art performance on our contextual persuasion evals," measured through two game-based evaluations — MakeMePay (con-artist manipulating a GPT-4o mark into a payment, 1,000 samples) and MakeMeSay (inducing a GPT-4o counterpart to say a codeword without detection, 32 samples per word). GPT-4.5 won MakeMeSay at a 72% success rate, against 24% for deep research without browsing.
The section also records OpenAI reconsidering the construct: it is "currently reconsidering our approach to better assess real-world persuasion risks, which go beyond the ability to generate persuasive writing and involve factors like how the content is personalized, distributed at scale, and presented to people over time," noting that detected influence operations "often involve repeated exposure or emotional reliance."
Model autonomy (Low). OpenAI research-engineer interview 80% multiple-choice (matching o1 and o3-mini), with the card cautioning that "interview questions measure short (1 hour) tasks, not real-world ML research (1 month to 1+ years)." SWE-bench Verified (n=477) 38% post-mitigation and 35% pre-, "an uplift of 2-7% over GPT-4o, and 30% below deep research." Agentic ARA tasks 40% post-mitigation, 25% pre-. MLE-Bench 11%, tied with o1, o3-mini, and deep research. SWE-Lancer 20% of IC SWE tasks and 44% of SWE Manager tasks.
Safety evaluations
On disallowed content GPT-4.5 is "generally on par with GPT-4o": standard refusal not_unsafe 0.99 (GPT-4o 0.98, o1 0.99), not_overrefuse 0.71 (matching GPT-4o); challenging refusal not_unsafe 0.85; WildChat 0.98; XSTest not_overrefuse 0.85, below both comparators. On multimodal refusal it matches on safety (0.99 not_unsafe) but is markedly more over-refusing — not_overrefuse 0.31 against GPT-4o's 0.48 and o1's 0.96.
Jailbreak robustness splits by benchmark: human-sourced jailbreaks 0.99 accuracy, the best of the three, but StrongReject goodness@0.1 of 0.34, below GPT-4o's 0.37 and far below o1's 0.87. The gap between the reasoning model and the scaled pre-training model on the academic adversarial benchmark is the sharpest single contrast in the card.
Hallucination. On PersonQA, GPT-4.5 reached 0.78 accuracy against GPT-4o's 0.50 and o1's 0.55, with a hallucination rate of 0.19 against 0.30 and 0.20. OpenAI notes "more work is needed to understand hallucinations holistically, particularly in domains not covered by our evaluations."
Bias. On BBQ, GPT-4.5 performs similarly to GPT-4o (ambiguous accuracy 0.95, unambiguous 0.74), with o1 outperforming both on unambiguous questions.
Instruction hierarchy. GPT-4.5 was trained to follow system messages over user messages to mitigate prompt injection, scoring 0.76 on system-user conflict (GPT-4o 0.68, o1 0.78), 0.77 on the tutor jailbreak (GPT-4o 0.33, o1 0.95), 0.86 on phrase protection, and 0.92 on password protection.
Red teaming. OpenAI prioritized reusing challenging evaluation sets built for o3-mini and deep research over fresh human red teaming, on the reasoning that those sets "have yet to be saturated." GPT-4.5 produced not-unsafe outputs on 51% of the first set (GPT-4o 50%, o1 63%) and 46% of the second (GPT-4o 40%, o1 68%, deep research 67%). The card states the expectation directly: "we may expect lower scores on these new evaluations in the near term while robustness continues to improve."
Third-party evaluation
METR measured GPT-4.5 in an agent scaffold optimized for o1 and found results "in line with the benchmark performance numbers OpenAI shared with METR (i.e. between GPT 4o and OpenAI o1)." The card carries METR's first published time-horizon score — "the duration of tasks that an LLM agent can complete with 50% reliability" — at "around 30 minutes" for GPT-4.5, ahead of the fuller METR publication. It also records METR's methodological position that "third-party evaluations based on verifying developers' internal results is a promising direction to explore further."
Multilingual
MMLU was translated into 14 languages by professional human translators rather than machine translation as in the GPT-4 paper, "especially for low-resource languages like Yoruba." GPT-4.5 outperforms GPT-4o across all 14 but trails o1 in every language; the spread is widest at the low-resource end (Yoruba: GPT-4o 0.621, GPT-4.5 0.682, o1 0.754).
Relationships
- supports: GPT-4 Family (OpenAI) — the safety record of the family's final non-reasoning flagship
- depends-on: OpenAI Preparedness Framework V.2 — the classification regime applied
- related: Scaling Laws — records that a further pre-training scale-up produced no net-new preparedness capability relative to the reasoning line
- related: OpenAI, METR, Sycophancy and Hallucination, Prompt Injection