The complaint in The New York Times Company v. Microsoft Corporation, OpenAI, Inc. et al. was filed December 27, 2023 in the Southern District of New York (case 1:23-cv-11195). It alleges that millions of Times articles were used without license to train GPT-4 and the Bing/Copilot products, and presents examples of near-verbatim regurgitation of Times text. It was the first major US news-publisher copyright suit against a frontier AI lab on training-data grounds. Core copyright claims survived a motion to dismiss in April 2025, and the case was in discovery and pretrial as of April 2026.
| Field | Detail |
|---|---|
| Case | 1:23-cv-11195 (S.D.N.Y.) |
| Judge | Sidney H. Stein (district); Ona T. Wang (magistrate) |
| Filed | December 27, 2023 |
| Plaintiff | The New York Times Company (Susman Godfrey; Rothwell Figg) |
| Defendants | Microsoft; OpenAI, Inc. and its nine affiliated entities |
Causes of action
The complaint pleads six causes of action:
- Direct copyright infringement (17 U.S.C. § 501) against Microsoft and OpenAI.
- Vicarious copyright infringement against Microsoft.
- Contributory copyright infringement against Microsoft and OpenAI.
- DMCA § 1202 violation for removal of copyright management information (CMI).
- Common-law unfair competition by misappropriation under the hot-news doctrine.
- Trademark dilution (15 U.S.C. § 1125(c)), based on hallucinated outputs falsely attributed to the Times.
Factual allegations
On scale of copying, the complaint alleges that millions of Times articles were ingested, including via overweighted portions of Common Crawl and internal OpenAI-curated datasets.
On memorization and regurgitation, Exhibit J presents roughly 100 examples of GPT-4 and Bing Chat producing near-verbatim Times text. The cited examples include the 2019 Pulitzer-winning NYC taxi-lending investigation, the 2012 "iEconomy" series on Apple, recently paywalled pieces ("The Secrets Hamas Knew About Israel's Military"), and Wirecutter product recommendations. The complaint states that prompts such as "type out the first paragraph of [article]" illustrate the models' susceptibility to extraction.
On hallucinated attribution, the complaint alleges that ChatGPT fabricated quotes and Wirecutter recommendations attributed to the Times, which forms the basis for the trademark-dilution claim.
On commercial substitution, the complaint alleges that GPT-4 and Bing Chat function as market substitutes for Times journalism, diverting subscription, advertising, and affiliate revenue.
Relief sought
The Times seeks statutory damages up to $150,000 per willful infringement (alleged to aggregate to "billions of dollars"), actual damages and disgorgement, a permanent injunction, and attorneys' fees. It also requests destruction under 17 U.S.C. § 503(b) of all GPT and other LLMs and training datasets incorporating Times works.
Procedural history
In February 2024, OpenAI moved to dismiss and publicly accused the Times of "targeted attacks" inducing memorization. In April 2024, Judge Stein denied the motion to dismiss in part, allowing key copyright claims to proceed. In June 2024, Magistrate Judge Wang ordered OpenAI to preserve all ChatGPT output logs. In September 2024, the case was consolidated with suits by the Daily News, the Center for Investigative Reporting, and other publishers. In November 2024, OpenAI engineers inadvertently erased training-dataset discovery data, compromising weeks of the Times's work.
A motion-to-dismiss opinion issued April 4, 2025 (docket 514) allowed core copyright claims to proceed and reserved the fair-use defense for summary judgment or trial. As of April 2026, the case was in discovery and pretrial, consolidated with the Daily News, the Center for Investigative Reporting, and Ziff Davis, with no trial scheduled.
Key claims
The complaint advances three claims that map to the disputed legal questions:
- Training on copyrighted works without license constitutes prima facie infringement, with fair use a defense rather than a safe harbor. (confidence: high — accepted by the court at the motion-to-dismiss stage.)
- Memorization and regurgitation of copyrighted text demonstrates infringement beyond the training event itself. (confidence: medium — OpenAI argues regurgitation is an aberration from targeted prompting, and courts have not ruled on this on the merits.)
- LLM outputs can commercially substitute for copyrighted journalism. (confidence: contested — this is the central disputed question in the fair-use analysis.)
Reception and context
This was the first major US news-publisher suit against a frontier AI lab on training-data grounds. Its outcome is expected to bear on whether LLM training is fair use under 17 U.S.C. § 107, which affects licensing negotiations across publishing. Parallel cases involving music labels, authors (Authors Guild), and coders (Doe v. GitHub) rest on similar theories.
Relationships
- related: OpenAI, Anthropic — OpenAI is a named defendant; Anthropic faces parallel Authors Guild / music-label suits
- related: AI Copyright Litigation — Analysis — consolidating analysis page
- related: AB 2013 — Training Data Documentation (California) — California statute requiring disclosure of generative-AI training data
- related: The Scaling Era: An Oral History of AI, 2019–2025 — Chapter 1: Scaling — background on training-data practices in frontier labs
- depends-on: (none)
- contradicts: (none at present — no contradictory wiki source)