California Assembly Bill 2013, the Generative Artificial Intelligence Training Data Transparency Act, is a state disclosure statute requiring developers of generative AI systems made available to Californians to post a high-level summary of the datasets used to develop or train those systems. It was authored by Assemblymember Rebecca Bauer-Kahan, signed into law on September 28, 2024, and codified at California Civil Code § 3111, with the disclosure obligation taking effect January 1, 2026.
| Jurisdiction | California |
| Bill ID | AB 2013 (2024) |
| Author | Assemblymember Rebecca Bauer-Kahan |
| Codified at | California Civil Code § 3111 |
| Status | Signed into law September 28, 2024; disclosure obligation effective January 1, 2026 |
Scope and definitions
The Act applies to "developers" of generative AI systems made available to Californians. The obligation reaches generative AI systems released on or after January 1, 2022, with the required documentation due by January 1, 2026.
Unlike SB 53, AB 2013 has no training-compute or revenue threshold, so it reaches small and large developers alike. The Act exempts AI used solely for security or integrity purposes, for the operation of aircraft in the national airspace, or developed for national-security, military, or defense purposes and made available only to a federal entity.
Key provisions
AB 2013 requires a developer to post on its website a high-level summary of the datasets used to develop or train the generative AI system. The required documentation must cover:
- Dataset sources and ownership — the sources or owners of the datasets.
- Purpose — how the datasets further the AI system's intended purpose.
- Scale — the number of data points in the datasets, described in general ranges.
- Data types — the types of data points in the datasets.
- Intellectual-property status — whether the datasets include data protected by copyright, trademark, or patent, or are in the public domain.
- Acquisition method — whether datasets were purchased or licensed by the developer.
- Personal information — whether the datasets contain personal information or aggregate consumer information as defined under California privacy law.
- Cleaning, modification, and timing — whether the developer cleaned, processed, or modified the datasets, and the dates of collection and first use.
AB 2013 is a disclosure statute rather than a restriction: it requires developers to document the intellectual-property status of their training corpora, but does not itself resolve whether training on copyrighted material is lawful. Closed-weight labs have historically disclosed little about training data and open-weight developers vary, so the § 3111 requirement establishes a partial transparency floor across both.
Interaction with copyright litigation
The required intellectual-property status disclosure intersects with AI training-data copyright suits, including Bartz v. Anthropic, Kadrey v. Meta, and Hachette et al. v. Meta (and Mark Zuckerberg). A developer's own § 3111 disclosure can become evidence in infringement litigation.
Anthropic's public AB 2013 compliance disclosure for the Claude model family is a primary document on frontier-lab training-data practice. It is summarized at the AB 2013 source page, which sets out its five data-source categories and its acknowledgment that the corpus includes third-party intellectual property "consistent with standard industry practice."
Constitutional challenge
The statute is the subject of a First Amendment challenge, SpaceXAI v. Bonta (AB 2013 challenge), filed by SpaceXAI on December 29, 2025 on compelled-speech and viewpoint-discrimination theories; a preliminary injunction was denied, and on July 23, 2026 Legal Advocates for Safe Science and Technology filed an amicus brief opposing the suit with nearly 30 co-signatories. An appeals ruling adopting strict scrutiny could bear on transparency provisions in SB 53, Illinois SB 315, and New York's RAISE Act (Source: transformernews.ai).
California's layered state AI regime
AB 2013 is one layer in California's stacked approach to AI governance, operating on the axis of training-data lineage:
| Statute | Focus | |
|---|---|---|
| AB 2013 | Training-data lineage transparency | |
| [[legislation/california-sb-53 | SB 53]] | Frontier-model safety disclosure |
| CCPA / ADMT regulations | Personal-information rights, automated-decision opt-out |
The three operate on different axes — data lineage, safety, and privacy — and a large California developer is subject to all three.
Relationships
- instance-of: AI Content Licensing — training-data governance and provenance.
- related: AB 2013 — Training Data Documentation (California) — source summary of the bill text and Anthropic's compliance disclosure.
- related: California SB 53 — sibling California transparency statute (safety axis).
- related: Bartz v. Anthropic, Kadrey v. Meta — training-data copyright litigation that AB 2013 disclosures intersect.
- regulated-by: California Attorney General.
Sources
- AB 2013 — Training Data Documentation (California) — source summary: AB 2013 bill text (Civil Code § 3111) and Anthropic's Claude-family compliance disclosure.