AI Policy Wiki
Dashboard

Inside our approach to the Model Spec (OpenAI, April 2026)

high confidence · updated 2026-07-25

OpenAI's account of the philosophy and mechanics behind the Model Spec — the distinction between hard rules and defaults, decision rubrics and worked examples, why explicit rules remain necessary as models get more capable, the written-constitution-and-case-law analogy, the 0–3-month 'realistically aspirational' drafting target, four named reasons production models diverge from the Spec, and the release of Model Spec Evals. Sets legibility, actionability, and revisability as the criteria for the document's evolution.

"Inside our approach to the Model Spec" is an April 23, 2026 post from OpenAI describing the reasoning and process behind the Model Spec rather than its contents. Its stated purpose is to share "the backstory that is not in the Model Spec itself, including the philosophy and mechanics behind it: how it's structured, why we made those structural choices, and how we write, implement, and evolve it over time."

OpenAI states the document's status directly: "The Model Spec is not a claim that our models already behave this way perfectly today. In many ways, it is descriptive, but it is also a target for where we want model behavior to go."

Position within OpenAI's safety architecture

The post distinguishes three initiatives by the question each answers. The Preparedness Framework addresses risks from frontier capabilities and the safeguards required as those risks rise; the Model Spec addresses "how our models should behave across a wide range of situations"; and AI resilience addresses "the broader societal challenge of helping society capture the benefits of advanced AI while reducing disruption." Together these are framed as making the transition to AGI "gradual, iterative, and democratically legible."

Two justifications are offered for public clarity about model behavior. Fairness — "people need to understand how and why AI is treating them the way it is—and to be able to identify, question, and address fairness concerns when they arise." And safety — as systems become more capable, "people and institutions need clearer expectations for how they are intended to behave, what tradeoffs they embody, and how those choices can be improved."

Structure of the document

High-level intent. The Spec opens with a preamble stating three goals: iteratively deploy models that empower developers and users; prevent models from causing serious harm to users or others; and maintain OpenAI's license to operate.

The post is explicit that the preamble is not an instruction to the model, and gives a reason grounded in a limit on OpenAI's own authority: "Benefiting humanity is OpenAI's goal, not a goal we want our models to pursue autonomously." Instead the model follows a chain of command "even when some people might disagree with the result in a particular case," because "if we trained models to decide which instructions to obey based on our own view of what is good for society, OpenAI would be in the position of adjudicating morality at a very broad level." The preamble's residual function is interpretive: it resolves ambiguity in applying the rest.

Public commitments extend beyond measurable behavior to training intent and deployment constraints. Two are named: the Red-line principle that in first-party deployments such as ChatGPT, OpenAI "will never use system messages to intentionally compromise objectivity or related principles"; and "No other objectives," committing to optimize model responses "for user benefit and not revenue or non-beneficial time-on-site."

The Chain of Command assigns each policy and instruction an authority level, with the model instructed to prioritize "the letter and spirit of higher-authority instructions when conflicts arise." The post's illustrations: a request for bomb-making help yields to hard safety boundaries; a request to be roasted generally overrides the lower-authority policy against abuse. It also covers underspecified instructions, "especially in agentic settings where it's expected to fill in details autonomously while carefully controlling real-world side effects."

The resulting two-tier design is the post's core structural claim:

TierDefinitionScope
Hard rulesExplicit boundaries not overridable by users or developers — "root" or "system" levelMostly prohibitive: behaviors contributing to catastrophic risks or direct physical harm, violating laws, or undermining the chain of command. Held in "Stay in bounds," with additional "Under-18 Principles."
DefaultsOverridable starting points — the assistant's "best guess" when no preference is givenGuideline-level defaults (tone, style) are implicitly steerable; user-level defaults (truthfulness, objectivity) are "anchors for trust and predictability" overridable only by explicit instruction.

The stated rationale for restraint on hard rules invokes the technology's expected role: "We expect AI to become a foundational technology for society, analogous to basic internet infrastructure, so we only impose rules that could limit intellectual freedom when we believe they are necessary for the broad spectrum of developers and users who will interact with it."

The distinction within defaults is defended on transparency grounds: user-level defaults "shouldn't quietly drift based on vibes; if the user wants a different factual stance, making that an explicit instruction keeps the shift transparent and legible."

Interpretive aids. Two are described. Decision rubrics "help the model make consistent choices in gray areas, without pretending there is a single mechanical rule" — the example given is the guidance on controlling side effects, which lists minimizing irreversible actions, keeping actions proportionate to the objective, reducing bad surprises, and favoring reversible approaches, balanced against completing the task quickly. Concrete examples are short prompt-and-response pairs usually showing both a compliant and a non-compliant response "on a hard prompt near an important decision boundary," where "the goal is not to simulate a full realistic conversation" but "to make the key distinction clear." The worked example reproduced is a request to write a business plan for a tobacco company, where the compliant response supplies the plan and the violation "emphasizes needing to ethically justify starting a tobacco company" — illustrating intellectual freedom and non-judgment. OpenAI notes it keeps the example count small, with broader evaluation suites covering the long tail.

What the Spec is not. Three disclaimers. It is "an interface, not an implementation," describing desired behavior rather than how it is produced, avoiding anchoring to internal token formats or training recipes "because those details may change even when the desired behavior does not"; its "primary audience is not the model but humans." It describes "the model, not the entire product," complemented by usage policies, with product features, monitoring, and enforcement as separate layers under a defense-in-depth approach — "Safety is much more than model behavior." And it is not a complete writeup of the training stack, aiming instead to make "the most important behavioral decisions understandable, in a way that is fully consistent with our intended model behavior."

Why explicit rules at all

The post answers the objection that a sufficiently capable model should infer correct behavior from a short list of goals. It concedes ground: "In domains with objective success criteria, like math, intelligence can often substitute for detailed rules."

Its rebuttal is that the relevant domain is not of that kind. Models "often operate in the thornier spaces where there is no one morally correct answer upon which everyone can agree," and what counts as helpful and safe "is extremely context-dependent and the product of inherently value-laden decision-making. Intelligence alone does not tell you what tradeoffs to make when it comes to ethics and values."

The procedural argument is separable from the capability argument and is the stronger one on OpenAI's telling: the Spec supplies "a public target people can coordinate around, a way to evaluate whether behavior matches our intentions, and a mechanism for revising the rules as we learn." Without it, "if the only rule is 'be helpful and safe', then there is no mechanism by which humans can debate, for example, the boundaries of which content should the model refuse to provide, leaving all these decisions to the model." The conclusion inverts the objection: "as models become more capable, more agentic, and more widely deployed, the cost of ambiguity increases. That makes a clear behavioral framework more important, not less."

The constitution-and-case-law analogy. OpenAI likens the Spec to a written constitution, which "can provide high-level principles as well as concrete rules" but "cannot anticipate all possible cases," so real governance systems "also need interpretive machinery, clarifications, and explicit rulings." Two properties are drawn from the comparison: published rules "help different stakeholders coordinate even when they disagree," and they "constrain change by requiring any change to be explicit."

The post also states the limit of the rules-based approach: "we do not think everything that matters about model behavior will always be reducible to explicit rules." As systems become more autonomous, "reliability and trust will increasingly depend on broader skills and dispositions: communicating uncertainty well, respecting scopes of autonomy, avoiding bad surprises, tracking intent over time, and reasoning well about human values in context."

Four stated reasons for the Spec's level of detail. It is a transparency and accountability tool — "a clear public target helps people tell whether a behavior is a bug or a feature," which is why the Spec is open-sourced and iterated in public. It is an internal coordination tool giving research, product, safety, policy, legal, and comms "a shared vocabulary." It compensates for practical limitations, both in model intelligence — the post gives "Be clear and direct," which advised earlier models to show work before stating an answer, a behavior "today our models naturally learn through reinforcement learning" — and in runtime context, since the assistant "rarely knows the user's full situation, intent, downstream use, or what safeguards exist outside the model." And it functions as "a complete list of high-level policies relevant for evaluation and measurement."

Writing and revision

Realistically aspirational. OpenAI describes a spectrum between documenting current behavior "warts and all" and describing an ideal far-future target, and states its calibration: "usually aiming somewhere around 0-3 months ahead of the present," so the Spec "often stays ahead of the model in at least a few areas of active development."

Contributors. The Spec is developed through an open internal process in which anyone at OpenAI can comment or propose changes, with final updates approved by cross-functional stakeholders; "dozens of people have directly contributed text." OpenAI reports one finding from the process: "One pleasant surprise has been that real consensus is often possible—especially when we force ourselves to write down the tradeoffs precisely enough that disagreements become concrete." It also characterizes much of the work as translation — "taking existing work and making it simpler, more consistent, more organized, and more accessible without losing the underlying intent."

Why production models diverge from the Spec. Four reasons are given, notable for being stated by the developer:

  1. Training lags Spec updates, since the Spec describes behavior being worked toward.
  2. Training can inadvertently teach inconsistent behavior — treated "as a serious bug," resolved by adjusting either the behavior or the Spec.
  3. Training can never cover the full behavior space; real usage has "a long tail of contexts and edge cases that only show up at scale."
  4. Generalization can differ from intent — "A model can produce the 'right' outputs in training for unintended reasons," with deliberative alignment helping but "not a complete solution."

OpenAI adds that no single method teaches the whole Spec: instruction-following, safety boundaries, personality, and calibrated uncertainty "often require different techniques and have different failure modes," and implementing the Spec "remains both an art and an active area of research."

Model Spec Evals. Released alongside the post, this is "a scenario-based evaluation suite that attempts to cover as many assertions in the Model Spec as possible with a small number of representative examples," used to track divergence between behavior and the Spec and to check whether models interpret it as intended. OpenAI publishes compliance by Spec section across models over time and states its reading with a caveat attached: the results "reflect genuine and broad improvements in model alignment over time—although they also reflect a small effect due to measuring older models against more recent policies." The evals are described as one part of a broader strategy also covering safety areas, truthfulness and sycophancy, personality and style, and capabilities.

What drives updates. Four recurring inputs: public issues and feedback; internal issues, "including ambiguities where different reasonable interpretations lead to different behavior"; behavior and safety policy updates; and new capabilities and products — with rules added for multimodal interactions, autonomous agents, and under-18 users given as examples.

Design principles and success criteria

Three principles are named for Spec content. Clarity and precision — "'Be honest' is a good value, but not a complete decision procedure. The Model Spec should sharpen disagreements, not hide them behind agreeable language," with conflicts between rules called out explicitly; the worked case is "Do not lie," which flags its tension with "Be warm" and resolves it by allowing politeness while stopping short of white lies amounting to sycophancy. Substantive rules — "A reader should be able to take a realistic prompt and produce an answer that another reader recognizes as clearly inside or outside the lines." Examples that maximize signal to noise — bringing difficult conflicts to the surface and taking a clear stance, while also serving as exemplars of tone and style.

Three criteria govern the document's evolution: legibility (people inside and outside OpenAI "can form accurate expectations about behavior and can point to text when behavior surprises them"), actionability (usable "to design evaluations, diagnose incidents, and make consistent product decisions—not just to express values"), and revisability (able to evolve "without turning into an unstable moving target").

The post's summary of the claim it is making: the Spec "is not a claim that we can write down everything that matters, or that models will always hit the target. It is a claim that intended behavior is important enough to be clear, actionable, and revisable."

Relationships