AI Policy Wiki
Dashboard

California AB 2013 — Generative AI Training Data Transparency

medium confidence · updated 2026-07-23

California law requiring developers of generative AI systems to publicly post documentation of the data used to train those systems — sources, ownership, IP status, personal-information handling, and acquisition methods. Codified at California Civil Code Section 3111; disclosure obligation effective January 1, 2026.

California Assembly Bill 2013, the Generative Artificial Intelligence Training Data Transparency Act, is a state disclosure statute requiring developers of generative AI systems made available to Californians to post a high-level summary of the datasets used to develop or train those systems. It was authored by Assemblymember Rebecca Bauer-Kahan, signed into law on September 28, 2024, and codified at California Civil Code § 3111, with the disclosure obligation taking effect January 1, 2026.

JurisdictionCalifornia
Bill IDAB 2013 (2024)
AuthorAssemblymember Rebecca Bauer-Kahan
Codified atCalifornia Civil Code § 3111
StatusSigned into law September 28, 2024; disclosure obligation effective January 1, 2026

Scope and definitions

The Act applies to "developers" of generative AI systems made available to Californians. The obligation reaches generative AI systems released on or after January 1, 2022, with the required documentation due by January 1, 2026.

Unlike SB 53, AB 2013 has no training-compute or revenue threshold, so it reaches small and large developers alike. The Act exempts AI used solely for security or integrity purposes, for the operation of aircraft in the national airspace, or developed for national-security, military, or defense purposes and made available only to a federal entity.

Key provisions

AB 2013 requires a developer to post on its website a high-level summary of the datasets used to develop or train the generative AI system. The required documentation must cover:

  1. Dataset sources and ownership — the sources or owners of the datasets.
  2. Purpose — how the datasets further the AI system's intended purpose.
  3. Scale — the number of data points in the datasets, described in general ranges.
  4. Data types — the types of data points in the datasets.
  5. Intellectual-property status — whether the datasets include data protected by copyright, trademark, or patent, or are in the public domain.
  6. Acquisition method — whether datasets were purchased or licensed by the developer.
  7. Personal information — whether the datasets contain personal information or aggregate consumer information as defined under California privacy law.
  8. Cleaning, modification, and timing — whether the developer cleaned, processed, or modified the datasets, and the dates of collection and first use.

AB 2013 is a disclosure statute rather than a restriction: it requires developers to document the intellectual-property status of their training corpora, but does not itself resolve whether training on copyrighted material is lawful. Closed-weight labs have historically disclosed little about training data and open-weight developers vary, so the § 3111 requirement establishes a partial transparency floor across both.

The required intellectual-property status disclosure intersects with AI training-data copyright suits, including Bartz v. Anthropic, Kadrey v. Meta, and Hachette et al. v. Meta (and Mark Zuckerberg). A developer's own § 3111 disclosure can become evidence in infringement litigation.

Anthropic's public AB 2013 compliance disclosure for the Claude model family is a primary document on frontier-lab training-data practice. It is summarized at the AB 2013 source page, which sets out its five data-source categories and its acknowledgment that the corpus includes third-party intellectual property "consistent with standard industry practice."

Constitutional challenge

The statute is the subject of a First Amendment challenge, SpaceXAI v. Bonta (AB 2013 challenge), filed by SpaceXAI on December 29, 2025 on compelled-speech and viewpoint-discrimination theories; a preliminary injunction was denied, and on July 23, 2026 Legal Advocates for Safe Science and Technology filed an amicus brief opposing the suit with nearly 30 co-signatories. An appeals ruling adopting strict scrutiny could bear on transparency provisions in SB 53, Illinois SB 315, and New York's RAISE Act (Source: transformernews.ai).

California's layered state AI regime

AB 2013 is one layer in California's stacked approach to AI governance, operating on the axis of training-data lineage:

StatuteFocus
AB 2013Training-data lineage transparency
[[legislation/california-sb-53SB 53]]Frontier-model safety disclosure
CCPA / ADMT regulationsPersonal-information rights, automated-decision opt-out

The three operate on different axes — data lineage, safety, and privacy — and a large California developer is subject to all three.

Relationships

Sources