Author: Nita Farahany Source: https://nitafarahany.substack.com/p/ailawpolicyclass5 Published: September 9, 2025
The fifth installment of Nita Farahany's "Inside My AI Law & Policy" course (Class 5 of 27), published September 9, 2025, covers training data and copyright. It is built around two reference points: the Anthropic author class settlement of $1.5 billion, announced September 5, 2025 and described as the largest copyright settlement in US history, and the NYT v. OpenAI complaint. The class walks through the five technical stages by which text becomes training data, then applies the four-factor fair use test to argue that each stage involves copying and that copyright doctrine collides with how AI models are trained.
How text becomes training data
The class frames the scale of training through a reading comparison: GPT-4 was trained on roughly 13 trillion tokens (about 10 trillion words). At a reading rate of 250 words per minute, that corpus would take about 76,000 years of continuous reading; the model's training run took 90–100 days.
Farahany breaks the process into five stages, arguing that copying occurs at each:
- Download — every download creates a copy. nytimes.com was a top-15 source for GPT-3 by volume.
- Tokenization — text is broken into subword units; common words map to single tokens, while a rare word such as "Pneumonoultramicroscopicsilicovolcanoconiosis" becomes 14 tokens.
- Training loop — masked language modeling, in which the model predicts a hidden word ("The cat [MASK] on the mat" → predict "sat") across billions of iterations.
- Fine-tuning — described as comparatively easy, since a few thousand examples can adapt model behavior, and the same technique can be used to remove safety features.
- Deployment — models sometimes reproduce training data verbatim, a behavior known as memorization, which OpenAI characterizes as a bug.
Fair use four-factor analysis
The class applies the four statutory fair use factors to AI training:
- Transformative use — AI converts text into mathematical representations, and the Google Books precedent for scanning supports a transformative reading, but model output often competes directly with the originals.
- Nature of the work — AI training deliberately favors high-quality and creative content, producing the paradox that the most protected works are also the most useful for training.
- Amount used — training usually ingests the entirety of a work, whereas earlier fair use cases involved limited copying.
- Market effect — AI-generated summaries are associated with news traffic drops of 20–30%, posing threats to investigative journalism, and an emerging licensing market may collapse before it matures.
The Anthropic settlement and the 2019 data choice
Farahany works through the arithmetic of the Anthropic $1.5 billion settlement as roughly 500,000 authors multiplied by $3,000 per work. The class notes that Judge Alsup questioned where the $3,000 figure came from, characterizing it as probably large enough to make the matter go away but small enough for the company to afford, against an authors' demand of $150,000 each.
The settlement is tied to testimony from Anthropic co-founder Ben Mann, who admitted downloading the Library Genesis dataset while at OpenAI in 2019 and said he "believed it was fair use at the time" — testimony the class presents as central to the lawsuit.
To illustrate the incentives facing an AI startup in 2019, the class lays out three options. Option A was to license everything legally, estimated at $2–5 billion, 18–24 months, and about 20% content coverage. Option B was to use shadow libraries, which was free, took about three days, reached roughly 80% coverage, and was illegal. Option C was to use public-domain works only, which excluded modern content. The class notes that a competitor took Option B and built a multi-trillion-dollar business.
Provenance
A Substack essay by Nita Farahany, published September 9, 2025 as Class 5 of her 27-part "Inside My AI Law & Policy" course. The class is anchored on the Anthropic author class settlement (announced September 5, 2025) and the NYT v. OpenAI complaint.
Relationships
- part-of: Nita Farahany intro course series (Class 5 of 27)
- related: AI Copyright Litigation — Analysis, NYT v. OpenAI and Microsoft — Complaint (Dec 2023), Training Data Governance (planned)
- previous: Inside My AI Law & Policy Class 4: The AI Control Paradox (Farahany, September 2025) next: Inside My AI Law & Policy Class 6: Training Data, Discovery Wars, and Who Gets Paid (Farahany, September 2025)