AI Policy Wiki
Dashboard

Hachette et al. v. Meta (and Mark Zuckerberg)

high confidence · updated 2026-06-06

Class-action copyright lawsuit filed May 5, 2026 in the Southern District of New York by Hachette, Macmillan, McGraw Hill, Elsevier, Cengage, and novelist Scott Turow against Meta Platforms and Mark Zuckerberg, alleging torrent-based shadow-library scraping for Llama training, contributory infringement, and DMCA §1202(b) CMI violations. Names Common Crawl as a source. Push U.S. AI copyright suit total to 105.

Hachette et al. v. Meta is a class-action copyright lawsuit filed May 5, 2026 in the U.S. District Court for the Southern District of New York by academic and trade publishers Hachette, Macmillan, McGraw Hill, Elsevier, and Cengage, together with novelist Scott Turow, against Meta Platforms and chief executive Mark Zuckerberg. The complaint alleges that Meta assembled training corpora for its Llama models from torrented shadow-library archives and seeks liability for direct and contributory copyright infringement and for violations of the Digital Millennium Copyright Act's copyright-management-information provisions. A May 7, 2026 filing named Common Crawl as a source of the allegedly copyrighted training data.

Infobox

FieldValue
PlaintiffsHachette Book Group, Macmillan Publishers, McGraw Hill, Elsevier, Cengage Learning, Scott Turow
DefendantsMeta Platforms, Inc.; Mark Zuckerberg
CourtU.S. District Court for the Southern District of New York
FiledMay 5, 2026
StatusActive (last checked May 10, 2026)

Background

The suit was brought by major academic and trade publishers (Hachette, Macmillan, McGraw Hill, Elsevier, Cengage) and novelist Scott Turow, who named Meta Platforms and CEO Mark Zuckerberg personally as defendants. According to one tally, the filing pushed the total count of U.S. AI copyright suits to 105 as of May 5, 2026 (Source: chatgptiseatingtheworld.substack.com).

A May 7, 2026 amended-complaint or related filing named Common Crawl as a source of Llama's allegedly copyrighted training data, described as the first class-action complaint to specifically target the Common Crawl pipeline rather than treat it as anonymous web-scrape provenance (Source: insideaipolicy.com).

Claims

The complaint advances the following allegations:

  • Torrent-based shadow-library scraping of LibGen, Z-Library, and similar shadow-library archives to assemble Meta's pre-training corpora for Llama models.
  • Direct copyright infringement on plaintiffs' books and academic journal articles.
  • Contributory infringement for distributing models trained on infringing data.
  • DMCA §1202(b) Copyright Management Information (CMI) violations, alleging Meta stripped or altered CMI on the works it ingested.
  • Common Crawl as named source: the May 7 filing alleges Llama training pipelines drew on Common Crawl–distributed dumps that themselves contained copyrighted material.

According to commentary on the case, the plaintiffs' theory is calibrated to the acquisition-side liability gap that the early-2025 fair-use rulings left open. In Bartz v. Anthropic, the court held Anthropic liable on shadow-library acquisition even after holding training itself a fair use; Hachette v. Meta is described as seeking to pin the same acquisition-side liability while sidestepping the training-use fair-use defense.

Defenses

Meta said the courts have "rightly found that training AI on copyrighted material can qualify as fair use," anchoring its public posture on the early-2025 wave of fair-use rulings, notably the Bartz v. Anthropic and Kadrey v. Meta rulings holding training use fair while treating shadow-library acquisition as a separate liability question (Source: iapp.org).

Current status

The case is active in the Southern District of New York, with the docket last checked May 10, 2026. The May 7 filing naming Common Crawl is the most recent development on record.

The suit has drawn attention as the first class-action to name Common Crawl as a source rather than treat it as a neutral web-scrape pipeline; commentary notes that if discovery were to confirm Common Crawl ingest of shadow-library content, frontier labs relying on Common Crawl could be implicated (Source: insideaipolicy.com). The complaint also names Mark Zuckerberg personally, a feature it shares with Musk v. Altman though on different (copyright) doctrine. Commentary has characterized the filing, as the 105th U.S. AI copyright suit, as a marker that post-Bartz/Kadrey AI copyright suits have moved from experimental class actions toward a routine litigation channel comparable to BitTorrent-era peer-to-peer infringement suits (Source: chatgptiseatingtheworld.substack.com).

Relationships

Sources