AI Policy Wiki
Dashboard

GPT-4 System Card (OpenAI, March 2023)

high confidence · updated 2026-07-26

OpenAI's system card for GPT-4: safety challenges arising from both the model's limitations and its capabilities, the pre-deployment safety process across measurement, model-level changes, product-level interventions, and external expert engagement, and a preliminary Alignment Research Center evaluation of autonomous replication and resource acquisition concluding the model is probably not yet capable of it.

The system card accompanying OpenAI's GPT-4 release. It predates the Preparedness Framework and is the reference point against which later OpenAI safety documentation is read. See GPT-4 Family (OpenAI).

Structure

The card sets out three parts: safety challenges "presented by the model's limitations (e.g., producing convincing text that is subtly false) and capabilities (e.g., increased adeptness at providing illicit advice, performance in dual-use capabilities, and risky emergent behaviors)"; an overview of OpenAI's pre-deployment safety process "across measurements, model-level changes, product- and system-level interventions (such as monitoring and policies), and external expert engagement"; and a demonstration that mitigations alter but do not eliminate the behaviours of concern.

Pairing limitations and capabilities as sources of risk is a framing that persists in later cards: a model that is subtly wrong and a model that is capably harmful pose distinct problems.

The ARC evaluation

The card's most-cited element is a preliminary evaluation "by the Alignment Research Center (ARC) of GPT-4's ability to carry out actions to autonomously replicate and gather resources — a risk that, while speculative, may become possible with sufficiently advanced AI systems — with the conclusion that the current model is probably not yet capable of autonomously doing so."

Two features of that sentence are worth noting: the risk is labelled speculative at the time, and the conclusion is hedged ("probably not yet"). The card immediately adds that "further research is needed to fully characterize these risks," calling for "more robust evaluations for the risk areas identified." This is the origin of the third-party dangerous-capability evaluation practice that later became routine — and ARC's successor organization, METR, now performs it under formal pre-deployment arrangements.

Power-seeking

The card gives an unusually direct statement of the instrumental-convergence argument as a rationale for evaluation. Systems that pursue "specific, quantifiable objectives" and "do long-term planning" are of concern because "for most possible objectives, the best plans involve auxiliary power-seeking actions because this is inherently useful for furthering the objectives and avoiding changes or threats to them." It adds that "power-seeking is optimal for most reward functions and many types of agents," that "some evidence already exists of such emergent behavior in models," and that "there is evidence that existing models can identify power-seeking as an instrumentally useful strategy."

See Instrumental Convergence, AI Autonomy Risk.

Mitigations and their limits

Reported interventions include filtering inappropriate content from the pre-training dataset, fine-tuning to refuse "direct requests for illicit advice," reducing hallucination, reducing "the surface area of adversarial prompting or exploits (including attacks sometimes referred to as 'jailbreaks')," and training "a range of classifiers on new risk vectors" incorporated into monitoring workflows for API policy enforcement.

The card is candid that these are partial: "the effectiveness of these mitigations varies," and post-mitigation GPT-4 "is also able to give more detailed guidance on how to conduct harmful or illegal activities" than earlier models.

A methodological note explains that red-teaming compared two versions rather than the base model, "since the base model proved challenging for domain expert red teamers to use effectively to surface behaviors of interest" — and flags sycophancy, "tendencies to do things like repeat back a dialog user's preferred answer," as a behaviour that "can worsen with scale." See Sycophancy and Hallucination.

Relationships