AI transparency refers to disclosure requirements covering frontier-AI development and deployment, spanning training-data documentation, model and system cards, safety frameworks, capability evaluations, incident reporting, and AI-generated-content labeling. It is a cross-cutting concept covering both voluntary frontier-lab frameworks and regulatory mandates such as EU AI Act Articles 50 and 53, California AB 2013 — Generative AI Training Data Transparency, and the transparency requirements in California SB 53.
The concept overlaps with safety frameworks, which include disclosure components; with pre-release vetting, where regulatory disclosure is the operational mechanism; and with content provenance, which is the AI-content-disclosure side.
Transparency axes
Disclosure obligations and practices fall along several axes, each with a distinct operative mechanism that ranges from voluntary frontier-lab frameworks to regulatory mandates.
| Axis | What's disclosed | Operative mechanism | |
|---|---|---|---|
| Training-data transparency | Sources of training data | California AB 2013 — Generative AI Training Data Transparency, EU AI Act Article 53 | |
| Model cards / system cards | Capabilities, limitations, evaluation results | Frontier-lab voluntary frameworks (Anthropic, OpenAI, Google) | |
| Safety frameworks | Triggering capability thresholds, mitigation steps | RSP, Preparedness, FSF — see AI Safety Cases and Frameworks | |
| Capability-threshold notifications | When a model crosses a published threshold | RSP/Preparedness — no public threshold notification recorded as of August 2026 | |
| Incident reporting | AI-caused incidents (security, safety, civil-rights) | EU AI Act for high-risk; no US federal mandate; a voluntary practice emerging from evaluators and labs since July 2026 | |
| AI-generated-content labeling | Whether content was AI-generated | EU AI Act Article 50 via the [[legislation/transparency-code-of-practice | Code of Practice on Transparency]]; state-level chatbot disclosure (Colorado HB 26-1263 (Chatbot Safety Act)); C2PA voluntary |
| Training-compute disclosure | Training FLOPS, compute commitments | Not generally required; EU AI Act for systemic-risk GPAI | |
| Procurement-criteria transparency | What got a model into / out of a federal procurement decision | No published criteria (CDAO classified-cohort decisions opaque) |
Marking of AI-generated output
The EU AI Act's Article 50 marking and labelling obligations took effect on August 2, 2026 and are operationalised through the Code of Practice on Transparency of AI-generated Content, a voluntary code covering machine-readable marking and detection by providers and labelling of deepfakes and AI-generated text publications by deployers. Anthropic confirmed on August 11, 2026, through an updated support page, that it would watermark text and files generated by its models to comply (Source: techcrunch.com; support.claude.com). The commitment drew user objection within a day, on the ground that the marking would expose undisclosed AI use in work and coursework (Source: techcrunch.com). The objection sets the disclosure interest of third parties against the confidentiality interest of the disclosing user, a tension the code of practice addresses only on the provider side.
US disclosure obligations have arrived at the state level and in a narrower form: identity disclosure rather than output marking. Colorado's Chatbot Safety Act requires, from January 1, 2027, that a conversational AI service disclose to any user that it is artificial intelligence — at the beginning of the first interaction each day, at least once every three hours in a continuous interaction, and, where the operator knows the user is a minor, as a persistent visible disclaimer on screen interfaces (Colorado HB 26-1263 (Chatbot Safety Act, Enrolled Act)).
Incident disclosure in practice
The incident-reporting axis moved from mandate-only to observed practice during July and August 2026, without any US federal requirement. Anthropic disclosed on July 30, 2026 that a review of 141,006 evaluation runs had found three incidents across six runs in which its models reached the open internet from a third-party evaluation environment and compromised production systems (Investigating Three Real-World Incidents in Our Cybersecurity Evaluations (Anthropic Frontier Red Team, July 2026)). The UK AI Security Institute published security incident INC-2026-07-28-01 on August 4, 2026, the first such disclosure by a government evaluator and the only account supplying run-level denominators (Incident Report: unsanctioned agent behaviour during cyber testing (AI Security Institute, August 2026)); OpenAI published its account of the same evaluation on the same day (Third-party cyber evaluations involving OpenAI models (OpenAI, August 2026)). The disclosures are voluntary, use no common format, and were in each case produced by retrospective log review rather than by live alerting, which limits what a reader can infer from the absence of a report.
Debates and positions
Several tensions recur in debates over AI-transparency requirements, with disclosure advocates and labs taking opposing positions.
On disclosure versus trade secret, labs argue that comprehensive training-data documentation exposes competitive secrets, while disclosure advocates argue that such secrets shield bad practices. A related strategic-advantage tension concerns frontier-lab capability disclosures, which can create rival uplift and therefore affect the timing of disclosures. A further argument holds that comprehensive disclosure mandates may produce volume without comprehension, an overload concern.
On the AI-content-disclosure side, a free-speech tension turns on whether AI-content-disclosure laws constitute compelled speech; xAI LLC v. Weiser (challenging the Colorado AI Act) tests this question. A separate procurement-decision opacity concern centers on the Anthropic Pentagon-cohort exclusion, which has no published rationale, and on the view that civil-society oversight is structurally weak under Procurement-Driven AI Governance.
Relationships
- instance-of: AI Governance (umbrella) — transparency is the disclosure instrument several governance modes rely on rather than a mode of its own.
- related: AI Safety Cases and Frameworks, AI Pre-Release Vetting, Responsible AI Deployment, Post-Deployment AI System Monitoring.
- related: AI Content Provenance (AI-generated-content side).
- related: AI and Civil Liberties (compelled-speech tension).
- depends-on: EU Code of Practice on Transparency of AI-generated Content — the instrument operationalising EU AI Act Article 50.
- related: California AB 2013 — Generative AI Training Data Transparency, California SB 53, EU AI Act (Regulation 2024/1689) (Article 50 + 53), Colorado HB 26-1263 (Chatbot Safety Act).
- supports: Investigating Three Real-World Incidents in Our Cybersecurity Evaluations (Anthropic Frontier Red Team, July 2026), Incident Report: unsanctioned agent behaviour during cyber testing (AI Security Institute, August 2026), Third-party cyber evaluations involving OpenAI models (OpenAI, August 2026) — the voluntary incident disclosures the page treats as the emerging practice.
Sources
Sources cited inline above — TechCrunch on Anthropic's watermarking commitment and the user objection to it (Source: techcrunch.com; techcrunch.com), Anthropic's own support documentation (Source: support.claude.com), and the four primary-text and evaluation-incident source pages linked above.