AI Policy Wiki
Dashboard

AI Copyright Litigation — Analysis

medium confidence · updated 2026-08-16

Consolidating analysis of major US copyright suits challenging frontier-AI labs' use of copyrighted works in training data. Anchored on NYT v. OpenAI & Microsoft (S.D.N.Y. 2023), with parallel cases by Daily News, CIR, Ziff Davis, Authors Guild, and others.

A body of US copyright litigation has developed since 2023 over frontier AI developers' use of copyrighted works as training data. The first merits rulings, issued in mid-2025, divided the question in two: training a large language model on lawfully acquired works has so far been treated as fair use, while the acquisition of pirated copies has not. That distinction emerged from Bartz v. Anthropic, where a June 2025 summary-judgment ruling was followed by a $1.5 billion settlement that the Authors Guild describes as the largest copyright settlement in US history (Source: authorsguild.org). The cases span several rightsholder clusters — news publishers, book authors, music labels, software developers, and visual artists — and turn on overlapping questions of fair use, data provenance, and market effect.

The fair-use split: Bartz v. Anthropic

Andrea Bartz, Charles Graeber, Kirk Wallace Johnson v. Anthropic PBC (N.D. Cal.) was filed August 19, 2024 before Judge William Alsup. It was the first AI-training copyright case to produce a merits ruling. Judge Araceli Martínez-Olguín granted final approval of the $1.5 billion settlement on July 20, 2026, cutting class counsel's requested $187.5 million fee to $101.6 million (a 3.75 lodestar multiplier against the requested 6.92) (Source: chatgptiseatingtheworld.substack.com). Full case tracking is at Bartz v. Anthropic.

In June 2025, on summary judgment, Judge Alsup held that using books to train an LLM is fair use when the copies were lawfully acquired, characterizing that use as strongly transformative. He denied fair use for Anthropic's downloading of pirated copies from the shadow libraries LibGen and PiLiMi, holding that acquiring infringing copies is not excused by a later transformative use. He set a December 1, 2025 trial on piracy liability and statutory damages (Source: authorsguild.org).

In July 2025 the court certified a class of rightsholders of books Anthropic downloaded from LibGen and PiLiMi that were timely registered with the Copyright Office and carry an ISBN or ASIN. The class was certified for the piracy claim only, not for the training-as-fair-use question, which as decided applies only to the three named plaintiffs.

In September 2025, facing statutory-damages exposure across roughly 500,000 qualifying titles — statutory damages run $750 to $150,000 per work — Anthropic settled for $1.5 billion. The settlement was preliminarily approved September 25, 2025, amounting to at least about $3,000 per title, funded in four installments through 2027 (Source: authorsguild.org). On administration, the claims deadline was March 30, 2026, and the final approval (fairness) hearing was set for May 14, 2026, with distributions calculated thereafter.

The final approval (fairness) hearing was held May 14, 2026 before presiding Judge Araceli Martínez-Olguín in San Francisco. At the 75-minute hearing, seven objectors were each given two minutes; plaintiffs' lead attorney Justin Nelson reported the settlement claims rate had risen from 91.3% to 92.77%, and there was no indication the roughly $3,000–$3,100-per-work payout would change. The judge's questions focused on attorneys' fees and the structure of the cost reserve rather than the merits of the objections, and she ordered Anthropic to file a supplemental brief by May 21 addressing why late opt-outs should not be honored (Source: publishersweekly.com; authorsalliance.org). As of June 2026 final approval remained pending, with observers expecting it to be granted (Source: clarkhill.com). Separately, on May 14, 2026, twenty-eight authors who opted out of the settlement filed a new copyright suit against Anthropic requesting a jury trial, arguing that class-action treatment lets AI companies extinguish high-value copyright claims cheaply (Source: law.com).

The ruling separates how the data was used (training, treated as fair use so far) from how it was obtained (piracy, not fair use). Observers note that this distinction shifts plaintiff strategy across the docket toward acquisition and provenance — shadow-library sourcing — rather than the training act itself, and that it establishes a concrete damages floor that makes settlement attractive to well-capitalized defendants.

Kadrey v. Meta

Days apart from Bartz, Judge Vince Chhabria reached a fair-use result for Meta in Kadrey v. Meta on narrower reasoning, ruling for Meta on the record the plaintiffs built while signaling that a better-developed market-dilution theory — that AI-generated output floods and devalues the market for the originals — could defeat fair use in a future case. Read together, Bartz and Kadrey leave training-as-fair-use provisionally favored but unsettled, with market effect (fair-use factor 4) and data provenance as the contested issues.

News-publisher cases: NYT v. Microsoft, OpenAI et al.

NYT v. Microsoft, OpenAI et al., No. 1:23-cv-11195 (S.D.N.Y.), was filed December 27, 2023 before Judge Sidney H. Stein. As of mid-2026 it is the leading news-publisher test of training-data fair use, distinct from the book-author line that Bartz settled. The full source summary is at NYT v. OpenAI and Microsoft — Complaint (Dec 2023) and full case tracking at NYT v. Microsoft, OpenAI et al..

The complaint pleads four sets of claims: direct, vicarious, and contributory copyright infringement (17 U.S.C. § 501); removal of copyright management information under DMCA § 1202; unfair competition by misappropriation under the hot-news doctrine; and trademark dilution arising from hallucinated attributions.

To demonstrate infringement, the complaint presents roughly 100 examples of near-verbatim regurgitation of Times articles by GPT-4 and Bing Chat (Exhibit J). The cited works include the Pulitzer-winning 2019 NYC taxi-lending investigation, the 2012 "iEconomy" series on Apple, paywalled pieces including "The Secrets Hamas Knew About Israel's Military," and Wirecutter content, with an asserted affiliate-revenue impact.

The procedural history is as follows. On April 4, 2024 the motion to dismiss was denied in part, allowing the core copyright claims to proceed. In June 2024, Magistrate Judge Wang ordered OpenAI to preserve all ChatGPT output logs. In September 2024 the case was consolidated with Daily News, the Center for Investigative Reporting, and Ziff Davis. In November 2024, OpenAI engineers inadvertently erased discovery data gathered by the Times. On April 4, 2025 a second motion-to-dismiss opinion (docket 514) retained the core claims and reserved the fair-use defense for summary judgment or trial.

Discovery escalated through 2026. On January 5, 2026, District Judge Sidney Stein affirmed Magistrate Judge Wang's order compelling OpenAI to produce 20 million anonymized ChatGPT logs — about 0.5% of preserved logs — across the consolidated MDL, over OpenAI's privacy objections, holding that logs bear on whether ChatGPT outputs compete with or substitute for copyrighted works (Source: news.bloomberglaw.com). The order was the first instance of a court compelling large-scale production of user-generated AI outputs, rather than training data alone, in an AI copyright case. Summary-judgment briefing closed April 2, 2026 (Source: ailawsuittracker.com). On July 9, 2026, a group of publisher plaintiffs led by the Times and the New York Daily News moved to sanction OpenAI, alleging it withheld and destroyed key evidence — including concealing for two years its ability to search ChatGPT logs and deleting logs despite a preservation order; OpenAI called the allegations "blatantly false" (Source: reuters.com; arstechnica.com). Full procedural tracking is at NYT v. Microsoft, OpenAI et al.. As of July 2026 the case remains in pretrial with the sanctions motion pending and no trial date set.

Docket size

A running tally maintained by the "ChatGPT Is Eating the World" blog counted 128 U.S. copyright suits against AI companies as of July 19, 2026, up 13 since mid-June; recent additions included Shakespeare v. Anthropic, Gilbert v. Anthropic, Hachette Book Group v. Google, and S.A. Jamendo suits against Nvidia and Suno (Source: chatgptiseatingtheworld.substack.com). The figure is a self-published aggregation rather than an official docket count.

The fair-use question also reached a non-US court: on July 24, 2026, the High Court of Delhi denied ANI Media's preliminary-injunction motion against OpenAI, holding in a 135-page ruling that using copyrighted works to train AI models — and storing copies for training — is prima facie fair dealing under Section 52(1)(a) of India's Copyright Act, citing the Bartz, Kadrey, and Google Books decisions; the case can still proceed to trial (ANI Media v. OpenAI (High Court of Delhi)) (Source: chatgptiseatingtheworld.substack.com; chatgptiseatingtheworld.com).

Parallel case families

The litigation falls into rightsholder clusters that share theories of liability:

ClusterLead casesTheory
News publishersNYT v. OpenAI/MS; Daily News v. OpenAI/MS; CIR v. OpenAI/MS; Ziff Davis v. OpenAITraining-data ingestion plus regurgitation
AuthorsAuthors Guild v. OpenAI; Silverman v. OpenAI; Kadrey v. Meta (Llama)Ingestion of copyrighted books
Music labelsUMG/Sony/Warner v. Anthropic (song-lyric regurgitation); [[litigation/umg-v-sunoRIAA v. Suno]]; RIAA v. Udio; [[litigation/sony-v-udioSony v. Udio II]] (July 21, 2026 — 30,117 recordings a judge barred from the original case)Lyric reproduction; generative-audio training
CodeDoe v. GitHub/OpenAI/Microsoft (Copilot)Unlicensed open-source code reproduction
Visual art[[litigation/andersen-v-stability-aiAndersen v. Stability AI]] (N.D. Cal., Jan 13, 2023 — in discovery); Getty v. Stability AI (UK; [[litigation/getty-v-stability-aiUS case]] voluntarily dismissed Aug 2025)Image-dataset ingestion
Books (OpenAI)[[litigation/tremblay-v-openaiTremblay v. OpenAI]] (N.D. Cal., Jun 28, 2023 — transferred to S.D.N.Y. MDL 1:25-md-03143 in Apr 2025)Book-training ingestion
Shareholder derivativeSEIU Pension Plan Master Trust v. Narayen and Hirschberger v. Narayen (Adobe); [[litigation/anderson-v-microsoftAnderson v. Microsoft]] (Jun 30, 2026); an action against NVIDIA officers including Jensen Huang (Jul 31, 2026); [[litigation/rosen-v-cookRosen v. Cook]] (Apple, Aug 14, 2026)Breach of fiduciary duty and proxy-disclosure violations for approving or concealing training-data exposure

The shareholder-derivative cluster

A cluster that emerged in 2026 sues corporate officers and directors rather than the company, recasting training-data exposure as a governance failure. The plaintiff is a shareholder suing derivatively on the company's behalf; the pleaded wrong is that directors and officers approved training on unlicensed material, exposed the company to copyright liability, or made proxy-statement misrepresentations concealing that risk. Five such actions had been filed across four companies by mid-August 2026, against Adobe, Microsoft, NVIDIA and Apple (Source: chatgptiseatingtheworld.com).

Two pleading strategies are visible within the cluster. Anderson v. Microsoft rests on disclosure, building its case on the company's proxy-statement representations. The Apple complaint pleads the underlying data practices directly — naming the Books3 corpus, the Panda-70M video corpus and the voice-model recordings — and adds an Illinois Biometric Information Privacy Act count alongside the copyright allegations, plus a waste-of-corporate-assets count. Both plead an Exchange Act §14(a) / Rule 14a-9 count against the directors.

The cluster's legal exposure is structurally distinct from the direct-infringement cases above: it does not require the plaintiff to win the fair-use question, only to show that the defendants breached a duty in accepting or concealing the risk of losing it. Whether that distinction holds is untested — no court had ruled on the merits of any of the five actions as of August 16, 2026.

Several disputed questions run across the docket:

  • Fair use of LLM training under 17 U.S.C. § 107. The central disputed question. Transformativeness, market substitution, and the weight of factor 4 (market effect) are actively being litigated.
  • Memorization and regurgitation. Whether memorization and regurgitation alter the fair-use analysis. OpenAI argues regurgitation is a rare product of adversarial prompting; plaintiffs argue it proves infringing copies are embedded in the model.
  • DMCA § 1202 liability. Removal of attribution metadata from training-time copies.
  • Secondary liability for downstream use. When users extract memorized content, who is liable.

Policy interaction

Several enacted and proposed measures intersect with the litigation. California AB 2013 requires generative-AI developers to disclose training-data sources, reducing the information asymmetry at the heart of training-data litigation; post-Bartz, it speaks directly to the provenance question the rulings made central. Article 53 of the EU AI Act requires GPAI providers to publish a "sufficiently detailed summary" of training content, which enables rightsholder enforcement in the EU. The America's AI Action Plan signals federal interest in a pro-training-fair-use posture, though it is not yet enacted.

A parallel licensing market has developed alongside the litigation, evidenced by settlements involving AP, Axel Springer, Financial Times, News Corp, Vox, Time, and Meredith.

Implications for frontier labs

Across the cases, several operational patterns recur. The discovery burden is substantial: preservation orders for ChatGPT logs imply multi-year litigation costs. Output filtering has become standard, with post-2024 frontier models generally including verbatim-extraction defenses. Licensing deals reduce exposure but do not eliminate it; the Authors Guild and music-label suits proceed despite unrelated licensing activity. Model-weights destruction remains a possible remedy: the Times's demand under 17 U.S.C. § 503(b) for destruction of GPT models and datasets is still pending.

Relationships

Sources