AI Policy Wiki
Dashboard

Inside My AI Law & Policy Class 6: Training Data, Discovery Wars, and Who Gets Paid (Farahany, September 2025)

medium confidence · updated 2026-06-06

Sixth class in Nita Farahany's AI Law & Policy series, on training-data governance and discovery. Centers on Magistrate Judge Ona T. Wang's May 13, 2025 NYT v. OpenAI preservation order requiring OpenAI to preserve every conversation from 400M users indefinitely; covers the 91.3% NMA trigram-overlap evidence, the irreversibility of training data, and five control-spectrum governance approaches (permission / collective / compulsory / safe-harbor / fair-use).

Author: Nita Farahany Source: https://nitafarahany.substack.com/p/training-data-discovery-wars-and Published: September 14, 2025

This is the sixth class (of 27) in Nita Farahany's AI Law & Policy course series, published as a Substack essay on September 14, 2025. It addresses training-data governance and litigation discovery, organized around a preservation order in the NYT v. OpenAI copyright case, the evidentiary record over memorization of news content, the technical irreversibility of trained models, and a spectrum of approaches to compensating content creators.

Discovery and the NYT v. OpenAI preservation order

The class is anchored on Magistrate Judge Ona T. Wang's May 13, 2025 preservation order in NYT v. OpenAI, which Farahany characterizes as "unprecedented in scope and implications": a court-ordered indefinite preservation of every ChatGPT conversation from 400M users, overriding the GDPR right-to-deletion and OpenAI's own 30-day deletion promises. The order was affirmed by District Judge Sidney Stein on June 26, 2025.

Farahany frames the order against ordinary discovery practice, in which courts limit discovery to documents directly related to the parties; she describes this order as instead reaching an entire user base. She cites privacy scholars who argue the order is fundamentally different from any prior discovery order.

The preservation regime applies unevenly across product tiers: ChatGPT Free, Plus, and Pro users have their conversations preserved, while ChatGPT Enterprise and Edu users and Zero Data Retention API users are exempt. Farahany describes this as privacy becoming a luxury good.

Evidence of memorization

For the underlying copyright claims, Farahany cites the News Media Alliance (NMA) submission to the Copyright Office, which reported that "optimized prompts" produced text with 91.3% three-word-sequence (trigram) overlap with a Boston Globe article. The submission also documents overrepresentation of news content in training corpora: news is 0.02% of the internet but 0.15% of OpenAI's C4 dataset (a 7.5x overrepresentation) and 1.97% of OpenWebText (roughly 100x).

Irreversibility of trained models

Farahany uses a saffron-soup analogy to describe irreversibility: once training data is "cooked" into a model's 1.76T parameters, it cannot be extracted. She notes the NMA's acknowledgment that "Precise unlearning may be computationally infeasible for very large models." Removing specific content would require either identifying which of the 1.76T parameters were influenced, which she describes as nearly impossible, or retraining from scratch, which she estimates at roughly $100M and about 6 months and compares to a small city's annual energy use.

Approaches to compensation

The class presents a five-approach governance spectrum, run as a "Smurf-arrangement" exercise ordered from maximum creator control to maximum AI freedom:

  1. Permission-based — every piece of content requires explicit permission; maximum creator control, described as impossible at scale.
  2. Collective licensing — an ASCAP-style central organization collects fees and distributes them proportionally, raising the question of how to measure proportional use.
  3. Compulsory licensing with opt-out — Congress sets a rate (for example, $0.001/1000 words); on such a rate the NYT might receive $43.80/year against its $978M in digital revenue.
  4. Safe harbor with standards — a YouTube Content ID-style regime granting safe harbor to those who implement detection, takedown, and revenue sharing.
  5. Fair use free-for-all — the approach Farahany attributes to Singapore and Japan, and toward which she says American courts are trending.

Farahany describes the resulting compensation outcomes as two-tiered. Large publishers can strike licensing deals: she cites the Associated Press receiving roughly $5M/year from OpenAI and a News Corp deal worth $250M over 5 years. Bloggers and academics, she argues, receive nothing, because the transaction costs of negotiation exceed any possible payment.

Relationships