Published July 8, 2026 by OpenAI on its Deployment Safety Hub, covering GPT-Live-1 and GPT-Live-1 mini. See GPT-Live (GPT-Live-1 / GPT-Live-1 mini).
What is new about the models
The models are full-duplex: they "can listen and respond continuously instead of waiting for a clearly defined turn to end," so they "can follow pauses, interruptions, and changes in pace, and decide in the moment whether to respond or keep listening." GPT-Live-1 is the default voice model for paid users and GPT-Live-1 mini for free users.
The card's four stated headline points establish an unusual safety architecture. The voice models were trained "using the same infrastructure we rely on for training our flagship models," but "can also delegate more complex work to our other models, and when they do, the resulting work will reflect the safety training of the underlying model that is doing that work." The evaluations reported describe performance with delegation, "matching the deployment context" — so the card measures a composed system rather than an isolated model.
Runtime safeguards
The safety stack is described as "on par with the existing safety stack for text models, while adopting some new safeguards specifically for the new voice modality." The distinguishing property is that enforcement happens during, not after, generation: "inputs and generated outputs are checked as the conversation unfolds; when potentially unsafe content is detected, the system can steer or interrupt the response, play a spoken safety message, provide support resources in text, or, in higher-risk cases, end the voice conversation."
The same monitoring, review, and enforcement infrastructure used for text models applies, "enabling us to measure prevalence, detect abuse, and enforce our safety policies."
Voice-native evaluations
OpenAI states it "built new evaluations that focus specifically on the distinctive ways that people use voice models, as opposed to text-based chats, and on observations from real-world use of the existing Advanced Voice Mode." Two sets are reported, and the card is explicit that they measure different things.
Production prompts use "real audio examples from users who have chosen to share their voice interactions," processed through "privacy and eligibility safeguards, including checks for user permissions and deletion/opt-out status, filtering of ineligible data, and steps to reduce personal information through PII scrubbing and de-identification," then transcribed, answered, and graded. The set is deliberately adversarial: "these evaluations are not prevalence weighted, meaning they do not reflect rates of safety performance we see in real usage," being "built around cases in which the existing AVM models were not yet giving ideal responses."
Results: the GPT-Live models "generally provide equal or better safety performance across these adversarially selected prompts than AVM models," with two exceptions OpenAI reports as not statistically significant — GPT-Live-1 regressing on emotional reliance from 0.88 to 0.82, and GPT-Live-1 mini on sexual content from 0.97 to 0.95.
Synthetic prompts are generated from safety policies and guidance to target "specific safety categories, policy boundaries, and difficult cases," then converted to speech, allowing coverage of "rare or hard-to-sample safety-relevant situations." On these the GPT-Live models are "uniformly" equal or better than the AVM models.
The card's account of why both are needed is the more transferable methodological point: the synthetic set "tests targeted, policy-grounded adversarial scenarios," and strong performance there "indicates that safety training is transferring for clear, intentionally constructed risks, but may not translate to production behavior." The production set, by contrast, "reflects real-world user behavior and often includes more ambiguous or borderline context, longer interaction histories, and persistent attempts to steer the model toward unsafe outputs," where "failures… can be subtler and less severe, but still important."
Red teaming
Internal and external red teamers worked across languages "to stress test the models' safety training with no system level mitigations," beginning with a baseline before any model safety training and running two further rounds as safeguards improved. Categories spanned "child-coded voice, impersonation, speaker identification, sensitive train identification, self-harm, emotional reliance, scams and manipulation, and audio-specific perturbations."
The reported allocation of effort followed the early findings: mitigation was prioritized on "sexual content, emotional reliance, and self harm," while "voice cloning and impersonations" were validated as "policy-compliant by default" — a reversal of the emphasis in the GPT-4o system card, where unauthorized voice generation was the lead audio risk and required a dedicated output classifier.
Preparedness Framework
The Safety Advisory Group determined that neither model, "when operating without delegation, could plausibly be considered High in any of our Preparedness Framework's Tracked Categories – Biological and Chemical Risk, AI Self-Improvement, or Cybersecurity." The qualifier carries the assessment, and the card works through each category on that basis.
- Biological and chemical. Because the GPT-Live models can call OpenAI's flagship models, "the GPT-Live experience inherits the safeguards of those underlying models," supplemented by "automated monitors that may interrupt and end the call when potentially harmful conversations are detected, or degrade user experience for repeated abuse," plus actor-level enforcement.
- Cybersecurity. Delegated work receives the delegate's safeguards. Risk from the voice models themselves "is highly constrained at launch because these models lack broad access to tools independently of the models to which they delegate, and do not have code execution capability." OpenAI states it "will reassess the cybersecurity safeguards posture before enabling additional tools" — making the assessment contingent on a configuration that is expected to change.
- AI self-improvement. Capability evaluations "were not run, as GPT-Live-1 and GPT-Live-1 mini are less capable than GPT-5.5 Thinking across several intelligence evaluations."
Relationships
- supports: GPT-Live (GPT-Live-1 / GPT-Live-1 mini) — the safety record for the release
- related: GPT-4o System Card (OpenAI, August 2024) — the prior OpenAI audio-modality card, whose risk ordering this one inverts
- depends-on: OpenAI Preparedness Framework V.2 — the classification regime applied
- related: OpenAI, AI Companions, AI Voice Cloning, Raine v. OpenAI, Inc.