Author: Nita Farahany Source: https://nitafarahany.substack.com/p/training-data-discovery-wars-and Published: September 14, 2025
This is the sixth class (of 27) in Nita Farahany's AI Law & Policy course series, published as a Substack essay on September 14, 2025. It addresses training-data governance and litigation discovery, organized around a preservation order in the NYT v. OpenAI copyright case, the evidentiary record over memorization of news content, the technical irreversibility of trained models, and a spectrum of approaches to compensating content creators.
Discovery and the NYT v. OpenAI preservation order
The class is anchored on Magistrate Judge Ona T. Wang's May 13, 2025 preservation order in NYT v. OpenAI, which Farahany characterizes as "unprecedented in scope and implications": a court-ordered indefinite preservation of every ChatGPT conversation from 400M users, overriding the GDPR right-to-deletion and OpenAI's own 30-day deletion promises. The order was affirmed by District Judge Sidney Stein on June 26, 2025.
Farahany frames the order against ordinary discovery practice, in which courts limit discovery to documents directly related to the parties; she describes this order as instead reaching an entire user base. She cites privacy scholars who argue the order is fundamentally different from any prior discovery order.
The preservation regime applies unevenly across product tiers: ChatGPT Free, Plus, and Pro users have their conversations preserved, while ChatGPT Enterprise and Edu users and Zero Data Retention API users are exempt. Farahany describes this as privacy becoming a luxury good.
Evidence of memorization
For the underlying copyright claims, Farahany cites the News Media Alliance (NMA) submission to the Copyright Office, which reported that "optimized prompts" produced text with 91.3% three-word-sequence (trigram) overlap with a Boston Globe article. The submission also documents overrepresentation of news content in training corpora: news is 0.02% of the internet but 0.15% of OpenAI's C4 dataset (a 7.5x overrepresentation) and 1.97% of OpenWebText (roughly 100x).
Irreversibility of trained models
Farahany uses a saffron-soup analogy to describe irreversibility: once training data is "cooked" into a model's 1.76T parameters, it cannot be extracted. She notes the NMA's acknowledgment that "Precise unlearning may be computationally infeasible for very large models." Removing specific content would require either identifying which of the 1.76T parameters were influenced, which she describes as nearly impossible, or retraining from scratch, which she estimates at roughly $100M and about 6 months and compares to a small city's annual energy use.
Approaches to compensation
The class presents a five-approach governance spectrum, run as a "Smurf-arrangement" exercise ordered from maximum creator control to maximum AI freedom:
- Permission-based — every piece of content requires explicit permission; maximum creator control, described as impossible at scale.
- Collective licensing — an ASCAP-style central organization collects fees and distributes them proportionally, raising the question of how to measure proportional use.
- Compulsory licensing with opt-out — Congress sets a rate (for example, $0.001/1000 words); on such a rate the NYT might receive $43.80/year against its $978M in digital revenue.
- Safe harbor with standards — a YouTube Content ID-style regime granting safe harbor to those who implement detection, takedown, and revenue sharing.
- Fair use free-for-all — the approach Farahany attributes to Singapore and Japan, and toward which she says American courts are trending.
Farahany describes the resulting compensation outcomes as two-tiered. Large publishers can strike licensing deals: she cites the Associated Press receiving roughly $5M/year from OpenAI and a News Corp deal worth $250M over 5 years. Bloggers and academics, she argues, receive nothing, because the transaction costs of negotiation exceed any possible payment.
Relationships
- part-of: Nita Farahany intro course series (Class 6 of 27)
- related: NYT v. OpenAI and Microsoft — Complaint (Dec 2023), AI Copyright Litigation — Analysis, Data Irreversibility (planned)
- previous: Inside My AI Law & Policy Class 5: The $1.5 Billion Question (Farahany, September 2025) next: Inside My AI Law & Policy Class 7: Why China Quit US Chips (Farahany, September 2025)