Hachette et al. v. Meta is a class-action copyright lawsuit filed May 5, 2026 in the U.S. District Court for the Southern District of New York by academic and trade publishers Hachette, Macmillan, McGraw Hill, Elsevier, and Cengage, together with novelist Scott Turow, against Meta Platforms and chief executive Mark Zuckerberg. The complaint alleges that Meta assembled training corpora for its Llama models from torrented shadow-library archives and seeks liability for direct and contributory copyright infringement and for violations of the Digital Millennium Copyright Act's copyright-management-information provisions. A May 7, 2026 filing named Common Crawl as a source of the allegedly copyrighted training data.
Infobox
| Field | Value |
|---|---|
| Plaintiffs | Hachette Book Group, Macmillan Publishers, McGraw Hill, Elsevier, Cengage Learning, Scott Turow |
| Defendants | Meta Platforms, Inc.; Mark Zuckerberg |
| Court | U.S. District Court for the Southern District of New York |
| Filed | May 5, 2026 |
| Status | Active (last checked May 10, 2026) |
Background
The suit was brought by major academic and trade publishers (Hachette, Macmillan, McGraw Hill, Elsevier, Cengage) and novelist Scott Turow, who named Meta Platforms and CEO Mark Zuckerberg personally as defendants. According to one tally, the filing pushed the total count of U.S. AI copyright suits to 105 as of May 5, 2026 (Source: chatgptiseatingtheworld.substack.com).
A May 7, 2026 amended-complaint or related filing named Common Crawl as a source of Llama's allegedly copyrighted training data, described as the first class-action complaint to specifically target the Common Crawl pipeline rather than treat it as anonymous web-scrape provenance (Source: insideaipolicy.com).
Claims
The complaint advances the following allegations:
- Torrent-based shadow-library scraping of LibGen, Z-Library, and similar shadow-library archives to assemble Meta's pre-training corpora for Llama models.
- Direct copyright infringement on plaintiffs' books and academic journal articles.
- Contributory infringement for distributing models trained on infringing data.
- DMCA §1202(b) Copyright Management Information (CMI) violations, alleging Meta stripped or altered CMI on the works it ingested.
- Common Crawl as named source: the May 7 filing alleges Llama training pipelines drew on Common Crawl–distributed dumps that themselves contained copyrighted material.
According to commentary on the case, the plaintiffs' theory is calibrated to the acquisition-side liability gap that the early-2025 fair-use rulings left open. In Bartz v. Anthropic, the court held Anthropic liable on shadow-library acquisition even after holding training itself a fair use; Hachette v. Meta is described as seeking to pin the same acquisition-side liability while sidestepping the training-use fair-use defense.
Defenses
Meta said the courts have "rightly found that training AI on copyrighted material can qualify as fair use," anchoring its public posture on the early-2025 wave of fair-use rulings, notably the Bartz v. Anthropic and Kadrey v. Meta rulings holding training use fair while treating shadow-library acquisition as a separate liability question (Source: iapp.org).
Current status
The case is active in the Southern District of New York, with the docket last checked May 10, 2026. The May 7 filing naming Common Crawl is the most recent development on record.
The suit has drawn attention as the first class-action to name Common Crawl as a source rather than treat it as a neutral web-scrape pipeline; commentary notes that if discovery were to confirm Common Crawl ingest of shadow-library content, frontier labs relying on Common Crawl could be implicated (Source: insideaipolicy.com). The complaint also names Mark Zuckerberg personally, a feature it shares with Musk v. Altman though on different (copyright) doctrine. Commentary has characterized the filing, as the 105th U.S. AI copyright suit, as a marker that post-Bartz/Kadrey AI copyright suits have moved from experimental class actions toward a routine litigation channel comparable to BitTorrent-era peer-to-peer infringement suits (Source: chatgptiseatingtheworld.substack.com).
Relationships
- litigates: Meta AI, Mark Zuckerberg.
- related: NYT v. Microsoft, OpenAI et al. — the parallel premium-publisher AI-training case.
- related: AI Copyright — concept-level frame; this case is the May 2026 marker for shadow-library and Common Crawl as a litigation target.
- related: AI Copyright Litigation — Analysis — comparative analysis page tracking AI copyright suits.
- depends-on: Data Provenance, C2PA, and Watermarking — provenance and DMCA §1202(b) intersect; the CMI-stripping claim depends on a working theory of CMI persistence through training pipelines.
Sources
- (Source: chatgptiseatingtheworld.substack.com) — ChatGPT is Eating the World, May 5, 2026.
- (Source: iapp.org) — IAPP Daily Dashboard, May 6, 2026.
- (Source: insideaipolicy.com) — Inside AI Policy, May 7, 2026.