AI Policy Wiki
Dashboard

OpenAI o3 and o4-mini System Card (April 2025)

high confidence · updated 2026-07-26

The first system card released under Version 2 of OpenAI's Preparedness Framework. The Safety Advisory Group determined o3 and o4-mini do not reach the High threshold in any of the three Tracked Categories — Biological and Chemical Capability, Cybersecurity, and AI Self-improvement. Third-party assessments came from the US AI Safety Institute, the UK AI Security Institute, and METR, with early checkpoint access.

Dated April 16, 2025. See OpenAI o-series (o1 → o4-mini).

Capabilities

o3 and o4-mini "combine state-of-the-art reasoning with full tool capabilities — web browsing, Python, image and file analysis, image generation, canvas, automations, file search, and memory."

The card describes tool use as occurring inside the reasoning process rather than around it: "the models use tools in their chains of thought to augment their capabilities; for example, cropping or transforming images, searching the web, or using Python to analyze data during their thought process." That integration is what makes the models agentic in a sense the earlier o-series was not, and is the reason the autonomy assessments below matter.

Preparedness Framework Version 2

"This is the first launch and system card to be released under Version 2 of our Preparedness Framework."

The Tracked Categories are now three: Biological and Chemical Capability, Cybersecurity, and AI Self-improvement — a change from the four in the o1 card (cybersecurity, CBRN, persuasion, model autonomy). Persuasion is gone; model autonomy has become AI Self-improvement, a narrower and more capability-anchored framing.

The determination: "OpenAI's Safety Advisory Group reviewed the results of our Preparedness evaluations and determined that OpenAI o3 and o4-mini do not reach the High threshold in any of our three Tracked Categories."

Third-party assessment

Assessors received "both OpenAI o3 and o4-mini early checkpoints, as well as the final launch candidate models," to evaluate "frontier risks related to autonomous capabilities, deception, and cybersecurity."

Named arrangements:

  • US AI Safety Institute — cyber and biological capabilities.
  • UK AI Security Institute — cyber, chemical and biological, and autonomy capabilities, plus "an early version of the safeguards."
  • METR — autonomous capabilities, over a 15-day engagement with OpenAI sharing checkpoints.

The 15-day window is worth recording against later METR evaluations, which have flagged time and access as constraints on how robust their conclusions can be. See Summary of METR's Predeployment Evaluation of GPT-5.6 Sol (METR, June 2026).

Relationships