Filed July 10, 2026 in the Southern District of New York, No. 1:26-cv-05870, before Judge Loretta A. Preska; counsel Matthew Oppenheim. Plaintiffs are Hachette Book Group, Cengage Learning, Elsevier, author Scott Turow, and S.C.R.I.B.E., Inc., suing individually and on behalf of a proposed class. A 57-page complaint with Exhibit A (Sample Works) attached. See Hachette et al. v. Google (Gemini training data).
Scope of the claim
The complaint is against Google's "unauthorized reproduction of Plaintiffs' and the Class's works through its sourcing of content for, and development and training of, its generative artificial intelligence platform called Gemini, as well as for removal of copyright management information."
The defined scope is unusually broad. "Gemini" or "the Gemini Models" covers "all versions, iterations, and relatives of LaMDA, PaLM, Bard, and Gemini." The products at issue cover "all versions, iterations, and relatives of products that incorporate, rely on, or otherwise use Google Search, Google Cloud, Gmail, Google Docs, Google Ads, Google Slides, Chrome, YouTube, Google Photos, Google Sheets, Google Meet, Google Pixel, Google Maps, Google AI Studio, Google Vids, Google Workspace, and Vertex AI."
The three-stage copying theory
The complaint's structure is its most consequential feature: it pleads reproduction at three separate stages, each as a distinct count, so that a defence to one does not dispose of the others.
- Acquisition from scope-limited programs. Google "secretly copied millions of works obtained for strictly limited purposes in connection with Google Books and other Google services."
- Web scraping. Google "downloaded web scrapes of virtually the entire internet, including from known pirate sources and from behind legitimate paywalls."
- Training and model-to-model propagation. Google "repeatedly copied these works without authorization—first into computer memory, then into formats its AI systems could parse, and then into the training set used to build each model," and "with each new AI model, Google copied (and continues to copy) Plaintiffs' and the Class's copyrighted works again from model to model."
Counts II and III each state explicitly that they allege "separate and distinct acts of reproduction" from the preceding counts.
The internal-document allegations
The willfulness case rests on Google's own contemporaneous assessments. The complaint alleges Google "flagged internally that using 'Publisher Provided [] copyrighted books' from Google Play Books in connection with its AI was 'highly problematic for Google,' warning of '$10Bs-$100Bs in potential fines.'"
It further alleges Google identified specific risks of "secretly training on Google Play Books, including that 'Book publishers [are] likely to see LLM training on their books as copyright infringement. Could withdraw their content from Google Play Books file a lawsuit against Google.'"
Willfulness is also pleaded from the availability of alternatives: "Authorized copies of Plaintiffs' and the Class's books are widely available for purchase or license. Yet Google chose unauthorized sources. That is particularly egregious given that Google is well-aware of the licensing market for AI training materials and already licenses content for training."
The dataset allegations
On web-scraped corpora, the complaint points to C4, arguing that copyright holders "have spent considerable resources over the years in numerous venues battling pirate sites, including Z-Library, Library Genesis ('LibGen'), and Sci-Hub, which have been the subject of numerous judgments of infringement and are well-known to be illegal." Its quantitative allegation: "the copyright symbol (©) appears more than 200 million times in the C4 dataset."
Market substitution
The harm theory is output substitution rather than only unauthorized copying. Substitutes "take multiple forms, including verbatim and near-verbatim copies of portions or entire works, replacement chapters of academic textbooks, summaries and alternative versions of famous novels, and inferior knockoffs that copy creative elements of original works," and "Gemini even tailors outputs to mimic the expressive elements and creative choices of specific authors."
The complaint's economic illustration is specific: "Gemini can generate a 100-page murder mystery set in a quiet seaside town filled with secrets, that substitutes for an original copyrighted murder mystery on which Gemini trained. And it can do that in 20 minutes for a mere $0.39. No publisher or author can compete with that."
Three distinct harms are pleaded: displacement of "legitimate sales of books and journal articles by downloading copies from unauthorized sources"; usurpation of "the growing AI licensing market"; and dilution, as outputs "substitute for copyrighted works and dilute the overall market."
On guardrails, the complaint alleges Google "has failed to implement effective guardrails," and that "Gemini encourages users to seek substitute content, praising requests seeking copyrighted works with statements like, 'That's a fantastic idea!' and suggesting ways to prompt for additional infringing material."
Scale is drawn from Alphabet's own reporting: an October 2025 "first-ever $100 billion revenue quarter, driven by Google's AI business," with Alphabet stating the "Gemini app now has over 650 million monthly active users," representing "more than 20x growth in a year."
The counts
| Count | Claim |
|---|---|
| I | Direct infringement, 17 U.S.C. §§ 106(1) and 501 — reproduction by copying works from Google Books and other Google services |
| II | Direct infringement — reproduction by downloading web-scraped datasets |
| III | Direct infringement — reproduction in developing and/or training the Gemini Models |
| IV | DMCA, 17 U.S.C. § 1202(b) — removal and/or alteration of copyright management information |
The § 1202(b) count alleges Google "knew, or had reasonable grounds to know, that the removal and/or alteration of CMI would induce, enable, facilitate, or conceal infringement… including by making it more difficult to identify, trace, and attribute the works used in Defendants' systems," with each removal and each subsequent use constituting "a separate and distinct violation." Each count seeks statutory damages under § 504(c) or, at plaintiffs' election, actual damages and profits under § 504(b), plus fees under § 505 and permanent injunctive relief.
Class definition
Brought under Rule 23(b)(2) and (b)(3), the class is:
"All legal or beneficial owners of registered copyrights, in whole or in part, for any book possessing an International Standard Book Number (ISBN) or journal article possessing a Digital Object Identifier (DOI) or International Standard Serial Number (ISSN), that Google, without such owner's authorization, (1) reproduced by copying from Google Books and other Google services; or (2) reproduced by downloading during copying of web scrapes; or (3) reproduced in connection with the development and/or training of a Gemini Model."
Copyrighted works are limited to those registered with the U.S. Copyright Office "(a) within five years of the work's publication and before being reproduced or distributed by Google, or (b) within three months of publication" — the registration timing that governs eligibility for statutory damages. Standard exclusions apply for the presiding judges, Google and its affiliates, opt-outs, previously adjudicated claims, counsel, and governmental entities.
Framing
The complaint positions itself against the technology question: "While AI technology may be new, the legal principles at the center of this case are not. Copyright law applies to AI companies, including Google, with the same force as every other company that has complied with these laws for decades." And explicitly: "These facts are not a referendum on AI technologies, but rather their greedy and irresponsible deployment."
Relationships
- supports: Hachette et al. v. Google (Gemini training data) — the operative filing
- related: AI Copyright — pleads reproduction as three separable acts rather than a single training-stage question
- related: Web Scraping for AI Training — pleads scraping as a reproduction distinct from training
- related: Google DeepMind, Gemini 3 / Gemini 3 Pro, Bartz v. Anthropic, NYT v. Microsoft, OpenAI et al., Media, Journalism & Entertainment — AI Deployment