A system card is a document published by an AI developer alongside a model release describing the model's capabilities, evaluation results, safety testing, observed behaviors, and deployment mitigations. The genre spans several labels — Anthropic and OpenAI publish "system cards" for flagship releases, OpenAI issued a "model card" for its gpt-oss open-weights family (gpt-oss-120b & gpt-oss-20b Model Card (OpenAI, August 2025)), and Meta publishes "safety and preparedness reports" for its Muse models (see Muse Spark (Meta Superintelligence Labs)) — but the documents share a common function: they are the developer's own public account of what a model can do and what its testing found. Andrew Clearwater characterizes them as sitting "somewhere between a technical research paper and a regulatory filing" (System Cards Are the Most Important AI Governance Document You're Not Reading Carefully Enough (Clearwater, February 2026)).
Origin and lineage
The format descends from the "model cards" framework proposed by Mitchell et al. in 2019, which recommended short standardized documents reporting a trained model's intended use, evaluation conditions, and performance across demographic and environmental factors (Source: arxiv.org). Meta introduced the "system card" label in 2022 as a resource explaining how a deployed AI system — as opposed to a single model — works, using its Instagram feed-ranking system as the first example (Source: ai.meta.com). OpenAI's GPT-4 System Card (March 2023) applied the label to a frontier language model release, documenting safety challenges observed in pre-deployment testing and the mitigations applied before launch (Source: cdn.openai.com).
Frontier system cards have since grown substantially in length and scope. The Claude Opus 4.6 card runs about 200 pages; the Claude Mythos Preview card runs 244 pages with a 58-page alignment risk update companion; Meta's Muse Spark safety report is 158 pages (System Cards Are the Most Important AI Governance Document You're Not Reading Carefully Enough (Clearwater, February 2026), summarized at System Card Due Diligence).
Content and function
Contemporary frontier system cards typically document benchmark and capability evaluations, dangerous-capability assessments under the developer's safety framework (such as Anthropic's Responsible Scaling Policy ASL designations or OpenAI's Preparedness Framework determinations), red-teaming and third-party evaluation results, observed agentic and deceptive behaviors, safeguard testing including cases where safeguards failed, and acknowledged limits of the developer's own knowledge (System Cards Are the Most Important AI Governance Document You're Not Reading Carefully Enough (Clearwater, February 2026)). Individual cards tracked as foundational sources include Claude Mythos Preview System Card, Claude Opus 4.6 System Card, Claude Opus 4.7 System Card, System Card: Claude Sonnet 4.5 (Anthropic, September 2025), Claude Sonnet 4.6 System Card, GPT-5.3-Codex System Card, GPT-5.4 Thinking System Card, GPT-5.5 System Card (OpenAI, April 2026), and gpt-oss-120b & gpt-oss-20b Model Card (OpenAI, August 2025) — the last the first documented application of OpenAI's Preparedness Framework to an open-weights release (gpt-oss-120b & gpt-oss-20b Model Card (OpenAI, August 2025)).
The documents are voluntary and self-reported: they are written by the developer's own technical teams, published without external mandate, and describe testing conducted under conditions where the model may recognize it is being evaluated. Clearwater argues this makes them simultaneously incomplete and the most granular public source available on a given model (System Cards Are the Most Important AI Governance Document You're Not Reading Carefully Enough (Clearwater, February 2026)).
Debates and positions
Clearwater's February 2026 essay argues that system cards should be treated as a central artifact for buyer-side vendor risk assessment rather than a marketing document, and that documented behaviors function as legal notice to deployers; that argument and its six-item review checklist are covered at System Card Due Diligence (System Cards Are the Most Important AI Governance Document You're Not Reading Carefully Enough (Clearwater, February 2026)). A recurring caveat in the wiki's source pages is that card conclusions are the developer's own assessment of its own model and benefit from independent corroboration (e.g., gpt-oss-120b & gpt-oss-20b Model Card (OpenAI, August 2025)).
Relation to policy
System cards are one of the transparency mechanisms referenced in voluntary international commitments: the Seoul Frontier AI Safety Commitments (2024) affirmed practices including publication of safety frameworks and model transparency, and safety-framework documents such as responsible scaling policies define the capability thresholds that system cards report against (see AI Transparency). Proposals for mandatory third-party assessment and pre-release review (see AI Pre-Release Vetting) would formalize evaluation reporting that system cards currently provide voluntarily.
Relationships
- related: System Card Due Diligence — the practitioner discipline of reading them
- related: AI Transparency, AI Pre-Release Vetting, Alignment Risk Update
- instance-of relationships (inverse): the individual system-card source pages listed above declare
instance-ofthis concept - depends-on: Seoul Frontier AI Safety Commitments (2024) — voluntary-commitment context for safety-framework publication