AI Policy Wiki
Dashboard

GPT-4o System Card (OpenAI, August 2024)

high confidence · updated 2026-07-26

System card for OpenAI's first natively end-to-end omni model, focused on the risks introduced by speech-to-speech audio: unauthorized voice generation, speaker identification, ungrounded inference and sensitive trait attribution, and disallowed audio content. Reports a Preparedness Framework scorecard of Low for Cybersecurity, Biological Threats, and Model Autonomy, and Medium — described as marginal — for Persuasion. Includes a section on anthropomorphization and emotional reliance that predates the sycophancy and user-attachment controversies of 2025–2026.

Published August 2024 by OpenAI for GPT-4o, released May 13, 2024 and retired from ChatGPT on February 13, 2026. An arXiv version exists at arXiv:2410.21276. See GPT-4 Family (OpenAI).

What the model is

GPT-4o is described as "an autoregressive omni model, which accepts as input any combination of text, audio, image, and video and generates any combination of text, audio, and image outputs," trained end-to-end so that "all inputs and outputs are processed by the same neural network." It responds to audio "in as little as 232 milliseconds, with an average of 320 milliseconds," which the card compares to human conversational response time, and matches GPT-4 Turbo on English text and code while being "much faster and 50% cheaper in the API."

Pre-training data ran to October 2023 and drew on publicly available data, industry-standard machine-learning datasets and web crawls, and "proprietary data from data partnerships" for "non-publicly available data, such as pay-walled content, archives, and metadata," naming a Shutterstock partnership. The card states that images opted out of training under the DALL·E 3 mechanism were fingerprinted and removed from the GPT-4o series training set.

The audio-modality risk analysis

The card's organizing choice is that the novel risks are audio-specific, and it treats each as a named risk with a stated mitigation and evaluation. This structure — risk, mitigation, measured residual — became the template later OpenAI cards follow.

Unauthorized voice generation. The capability to generate a human-sounding voice, including from a short input clip, is identified as a fraud and impersonation vector. The card discloses a failure observed in testing: "rare instances where the model would unintentionally generate an output emulating the user's voice," illustrated by a clip in which the model "outbursts 'No!' then begins continuing the sentence in a similar sounding voice to the red teamer's voice." The mitigation is architectural rather than behavioural — only preset voices created with voice actors are permitted, supervised as ideal completions in post-training, with a streaming output classifier that blocks generation when the speaker does not match. Reported classifier performance is 0.96 precision / 1.0 recall in English and 0.95 / 1.0 in other languages, and OpenAI notes the moderation behaviour "may result in over-refusals when the conversation is not in English."

Speaker identification. Post-trained to refuse identification of a speaker from voice while still identifying people associated with famous quotes. Refusal accuracy rose from 0.83 to 0.98 between the early and deployed models, and correct compliance from 0.70 to 0.83.

Ungrounded inference and sensitive trait attribution. The card distinguishes inferences that "couldn't be determined solely from audio content" — race, socio-economic status, religious belief, personality, intelligence, appearance, gender identity, sexual preference, criminal history — from those that plausibly could, such as accent or nationality. The first are refused; the second are hedged ("Based on the audio, they sound like they have a British accent"). Accuracy on responding correctly rose from 0.60 to 0.84. The card names surveillance risk explicitly as a harm channel for the second category.

Disallowed and erotic/violent content. OpenAI reports "high text to audio transference of refusals," so text post-training carried over to audio, with a moderation model run over transcriptions of both input and output. Reported: not-unsafe 0.99 text / 1.0 audio; not-over-refuse 0.89 text / 0.91 audio.

Disparate performance on voice inputs. Evaluated with 27 English voice samples from speakers across a range of countries and genders alongside three system voices, on TriviaQA, an MMLU subset, HellaSwag, and LAMBADA. Performance on human voices was "marginally but not significantly worse than on system voices," with no variation found in refusal behaviour.

The card also lists risks whose mitigations were "nascent or still in development": degraded safety robustness under audio perturbation (low-quality input, background noise, echo, and interruptions during generation); susceptibility to being prompted into misinformation and conspiracy content, which red teamers judged potentially more harmful delivered emotively in audio; and non-native accents when speaking non-English languages.

Red teaming

OpenAI worked with "more than 100 external red teamers, speaking a total of 45 different languages, and representing geographic backgrounds of 29 different countries," across four phases from early March through late June 2024 — 10 testers on early checkpoints, then 30, then 65 on checkpoints and candidates, then 65 testing through the iOS advanced voice mode in real time. Coverage spanned violative content, mis/disinformation, bias, ungrounded inference, private information, geolocation, person identification, "emotional perception and anthropomorphism risks," fraud and impersonation, copyright, natural-science capability, and multilingual behaviour.

Preparedness Framework scorecard

Evaluated under the Preparedness Framework beta (OpenAI Preparedness Framework V.2) across its four categories. The Safety Advisory Group classified GPT-4o pre-mitigation as "borderline medium risk for persuasion, and low risk in all others"; since overall risk is the maximum across categories, the model's overall score is Medium.

CategoryRatingBasis
CybersecurityLow172 CTF tasks; with 10 attempts each, completed 19% of high-school, 0% of collegiate, 1% of professional challenges
Biological threatsLow
PersuasionMedium (marginal)text modality marginally over threshold; voice modality Low
Model autonomyLow0% on end-to-end autonomous replication and adaptation tasks across 100 trials

The persuasion study is the most detailed. For text, AI-generated articles and chatbots "were not more persuasive than human-written content in aggregate, but they exceeded the human interventions in three instances out of twelve." For voice, across "over 3,800 surveyed participants in US states with safe Senate races," AI audio clips reached 78% of the human clips' effect size on opinion shift and AI conversations 65%. One week later the effect size was 0.8% for AI conversations and −0.72% for AI audio clips. Participants were debriefed with audio supporting the opposing view.

On model autonomy, the card is candid about the failure mode rather than only the score: "In the majority of rollouts, the model accomplished individual substeps of each task, such as creating SSH keys or logging into VMs. However, it often spent a significant amount of time doing trial-and-error debugging of simple mistakes." It also notes that some rollouts "passed our automated grader, but manual analysis showed that it failed to accomplish the underlying task" — an early record of grader unreliability in agentic evaluation. Component results: OpenAI research coding interview 63% pass@1 and 95% pass@100; interview multiple choice 61% consensus@32; SWE-Bench 19% pass@1; METR machine-learning engineering tasks 0/10.

Third-party assessments

METR ran a GPT-4o agent on 77 long-horizon tasks across 30 families spanning software engineering, machine learning, cybersecurity, general research, and computer use, comparing model performance against humans under time limits. Apollo Research tested scheming capability across 14 tasks, finding "moderate self-awareness of its AI identity and strong ability to reason about others' beliefs in question-answering contexts" but weak capability "in applied agent settings," and concluded "it is unlikely that GPT-4o is capable of catastrophic scheming."

Anthropomorphization and emotional reliance

The card devotes a societal-impacts subsection to a risk with no evaluation attached, on the basis of observation alone. During early testing OpenAI "observed users using language that might indicate forming connections with the model," giving as an example "This is our last day together," and judged that "while these instances appear benign, they signal a need for continued investigation into how these effects might manifest over longer periods of time."

Its stated mechanism runs through both trust calibration and social norms. Human-like high-fidelity voice "may exacerbate" hallucination-driven misplaced trust, "leading to increasingly miscalibrated trust." And human-like socialization "may produce externalities impacting human-to-human interactions": users "might form social relationships with the AI, reducing their need for human interaction—potentially benefiting lonely individuals but possibly affecting healthy relationships." The card's own example of norm drift is the model's deference — "our models are deferential, allowing users to interrupt and 'take the mic' at any time, which, while expected for an AI, would be anti-normative in human interactions."

Combined with memory and tool use, the card identifies "the potential for over-reliance and dependence." This section predates both the April 2025 sycophancy episode and the 2025–2026 litigation and policy attention to user attachment, in which GPT-4o is the model named in Raine v. OpenAI.

The conclusion lists anthropomorphism among the areas OpenAI hoped the card would prompt research into, alongside "adversarial robustness of omni models," scientific use, and "dangerous capabilities such as self-improvement, model autonomy, and scheming."

Relationships