Fair Learning is a 2021 Texas Law Review article by Mark A. Lemley and Bryan Casey arguing that machine-learning training on copyrighted material qualifies as transformative fair use under existing US copyright doctrine. Written before ChatGPT and GPT-4, it is the academic articulation of the fair-use-for-AI-training position that AI labs including OpenAI, Anthropic, and Meta later invoked in copyright litigation (Source: texaslawreview.org).
Authors: Mark A. Lemley (William H. Neukom Professor, Stanford Law; partner, Durie Tangri LLP) and Bryan Casey (Fellow, Center for Automotive Research at Stanford) Publication: Texas Law Review Released: March 20, 2021 Source: texaslawreview.org
Summary of argument
Lemley and Casey argue that machine-learning systems learn by exposure to existing creative work in a manner they liken to human learning, and that copyright law has not traditionally treated such exposure as infringement. The paper opens with OpenAI's MuseNet music-generation system as an illustrative example. It anticipates the copyright-litigation wave that arrived in 2023–2025, and argues at length that training falls within transformative fair use.
The argument proceeds in four moves:
- The peculiarity of the copyright question. ML training requires copying copyrighted works wholesale into training datasets, even though the resulting model does not typically reproduce them. The authors contrast this with traditional fair-use cases (parody, criticism, reverse engineering), which involve narrower copying.
- The transformative-use case for ML training. Training is purpose-transformative even when not expression-transformative, because the resulting model serves a purpose distinct from consuming the original.
- The market-harm analysis. ML systems do not substitute for the original works in their primary markets (for example, reading the books or watching the videos), so the market-harm factor weighs weakly against AI labs.
- The "learning" framing. The authors analogize ML training to a human reading and learning from a book, a framing they argue makes the policy case for fair use compelling if accepted.
The paper acknowledges hard cases, including memorization, near-verbatim reproduction, and output that competes directly with training-data sources, but argues these are exceptions that fair-use doctrine can handle without breaking the general rule.
Role in the AI copyright debate
The paper was published in March 2021, before ChatGPT (November 2022), before the NYT v. OpenAI complaint (December 2023), and before the AI fair-use lawsuits accumulated to roughly 19 by mid-2024 (per Reisner, The Hypocrisy at the Heart of the AI Industry — Alex Reisner (The Atlantic, March 2026)). It is the pre-suit academic articulation of the transformative-use position that AI labs later defended in court, and serves as the doctrinal anchor for the defendant-side fair-use theory in AI copyright litigation. Analyses of that litigation engage Lemley and Casey directly as the primary academic statement of the defendant theory, including AI Copyright Litigation — Analysis and the entity page Mark Lemley, for which this is Lemley's most-cited single piece of AI-policy writing.
Lemley's own role as an AI-industry lawyer at Durie Tangri is disclosed in the paper's footnotes. The disclosure bears on the paper's standing as independent academic analysis without vacating its argument.
Contesting positions
Several sources contest the Lemley/Casey framing:
- The Hypocrisy at the Heart of the AI Industry — Alex Reisner (Atlantic, March 2026): Reisner contests the fair-use framing, citing AI memorization research showing that models can reproduce near-verbatim training data. Reisner's documentary find (a Schmidt video and the Amodei memo) operates as evidence against the transformative-use framing even though it does not engage Lemley and Casey directly.
- NYT v. OpenAI / Microsoft: The NYT plaintiff theory contests the Lemley/Casey position, with documented evidence that GPT models reproduce NYT articles verbatim. The Lemley/Casey position is the doctrinal basis of OpenAI's defense; the NYT theory contests it.
- Anthropic and Alignment — Ben Thompson (Stratechery, March 2026): Dario Amodei's 2021 internal Anthropic memo (later unsealed) acknowledges that AI training could be "an increasingly extractive concentrator of wealth," a position structurally contrary to the transformative-use framing.
- Ed Newton-Rex / Fairly Trained: An industry counter-voice that disputes the Lemley/Casey position as a matter of policy preference, independent of legal doctrine.
- AI Snake Oil — Narayanan and Kapoor (2024) and other empirical AI-skepticism sources: Lemley and Casey represent the legal-policy wing of the broader "AI is just learning, like humans" framing, which Bender, Hanna, and others contest from a linguistic and sociological angle. The "learning like humans" framing is itself contested in Ai And Language Models.
Confidence note
Confidence is high for the paper's argument and authorial provenance. It is medium for the descriptive claim that ML systems "learn the same way humans do," a contested philosophical and cognitive-scientific claim rather than a settled fact. The paper is explicit about this uncertainty; downstream industry uses of the paper sometimes are not.
Relationships
- supports: AI Copyright Litigation — Analysis, Mark Lemley
- contradicts: The Hypocrisy at the Heart of the AI Industry — Alex Reisner (The Atlantic, March 2026) (Reisner's documentary evidence undermines the transformative-use claim); Ed Newton-Rex / Fairly Trained (policy disagreement); NYT v. Microsoft, OpenAI et al. (plaintiff theory)
- related: Anthropic and Alignment — Ben Thompson (Stratechery, March 2026) (Amodei 2021 internal memo as related doctrinal artifact), AI Snake Oil — Narayanan and Kapoor (2024) (broader AI-skepticism context), Ai And Language Models (the "learning like humans" framing is contested in this concept space)
- instance-of: AI Safety Cases and Frameworks (legal-academic articulation)