AI copyright covers the legal and economic questions about whether and how copyrighted works can be used to train AI systems, and what legal status the resulting outputs have. Specific US cases are set out at AI Copyright Litigation — Analysis; the doctrinal and policy frame underlying them is described below.
The questions grouped under "AI copyright" fall into four distinct disputes: ingestion liability (whether copying works into a training set is an infringement), output liability (whether a generated work that reproduces or substantially mimics a copyrighted input is an infringement), authorship and protectability (whether AI outputs themselves can receive copyright), and licensing architecture (if some uses require licenses, who grants them, on what terms, to whom).
Doctrinal axes by jurisdiction
United States: fair use, four factors, § 1202
The principal US defense is fair use under 17 U.S.C. § 107. Labs argue that training is transformative (factor 1, purpose), that it uses works as input to a statistical process rather than for their expressive content (factor 2, nature), that scale is justified by transformativeness (factor 3, amount), and that training does not substitute for the market for originals (factor 4, effect). Plaintiffs contest all four, particularly factor 4, where output markets such as news summaries, stylistic imitation, and functional substitution are argued as the cognizable harm.
DMCA § 1202 (copyright management information) gives plaintiffs a separate track: labs that stripped attribution metadata from training copies face statutory damages regardless of how fair use is resolved. This is the live theory in NYT v. OpenAI (NYT v. OpenAI and Microsoft — Complaint (Dec 2023)).
European Union: Article 4 TDM exception plus opt-out
The EU DSM Directive (2019) Article 4 creates a general text-and-data-mining (TDM) exception to copyright, under which rightsholders may opt out by machine-readable reservation. The EU AI Act Article 53 layers on top: GPAI providers must respect opt-outs, publish a "sufficiently detailed summary" of training content, and maintain a policy to comply with EU copyright law globally rather than only for EU training runs.
This is an opt-out regime with disclosure, distinct from US fair use and from the UK's unresolved approach. Opting out is a unilateral act, and the disclosure obligation creates the evidentiary basis for enforcement.
United Kingdom: the 2025 opt-out controversy
The UK's initial 2022 IPO proposal was a broad TDM exception with no opt-out, the most lab-friendly approach in any major jurisdiction. Opposition from creative industries, including UK Music, the Publishers Association, and the Society of Authors, led the Sunak government to pause that plan in February 2023.
The Starmer government's late-2024 consultation revived a TDM exception with opt-out modeled on the EU approach. The 2025 response, documented in the UK Data (Use and Access) Act 2025 debates, met a coordinated backlash including a Times open letter signed by Paul McCartney, Kate Bush, Thom Yorke, Stephen Fry, and more than 400 other creators, and a silent protest album released by 1,000 British musicians. The bill was amended multiple times in the Lords, and the final text defers the question to further consultation, leaving it unresolved.
Japan: Article 30-4 and the JCP experiment
Japan's 2018 revision of Article 30-4 of the Copyright Act created the broadest TDM exception of any major jurisdiction: any copyrighted work can be used for "non-enjoyment" purposes including machine learning, without permission and without opt-out, except where the use would "unreasonably prejudice the interests of the copyright owner." This made Japan a favored training-data jurisdiction for a period. METI clarifications (2024) tightened the "unreasonable prejudice" language in response to music-industry concerns, but Japan remains the most permissive G7 regime.
Indonesia: 2026 copyright rewrite
Indonesia's rewrite of its copyright law, reported July 17, 2026, was described as putting Google and AI platforms "on notice" over the use of copyrighted works, extending the training-data licensing debate to a major Southeast Asian jurisdiction (Source: reuters.com).
Case clusters
For detailed status, see AI Copyright Litigation — Analysis. Copyright lawsuits filed against AI companies reached 125 on July 4, 2026, with OpenAI the most-sued defendant and book authors the largest plaintiff group (Source: chatgptiseatingtheworld.substack.com).
News publishers. NYT v. OpenAI/Microsoft (S.D.N.Y., filed December 2023) is the central case. Parallel actions by Daily News, Center for Investigative Reporting, Ziff Davis, and Dow Jones / New York Post v. Perplexity extend the theory across the major US news-publisher universe. The settlements track a parallel licensing market: AP, Axel Springer, Financial Times, News Corp, Vox, Time, and Meredith have signed OpenAI licensing deals, while NYT specifically declined.
Authors. Authors Guild v. OpenAI (S.D.N.Y., 2023) and Silverman v. OpenAI lead the literary cluster. Kadrey v. Meta (N.D. Cal.) produced the first substantive merits ruling for either side: Judge Chhabria's March 2025 decision granted partial summary judgment for Meta on fair-use grounds on some claims (transformativeness) while leaving market-effect claims alive. Authors Guild briefs argue at length that market effect on author earnings is the decisive factor 4 question.
Visual art. *Andersen v. Stability AI* (N.D. Cal., filed January 13, 2023) and Getty Images v. Stability AI (UK High Court 2023; US parallel) litigate both ingestion and output-infringement theories. The Andersen case was narrowed but survives on key claims and remained in discovery as of July 2026. The US Getty action was voluntarily dismissed on August 14, 2025 and terminated four days later without a merits ruling, leaving the UK proceeding as the only Getty case to have generated substantive law (Source: courtlistener.com). Bartz v. Stability (2024) adds photographers' claims specifically on identifiable-style replication.
Music. UMG, Sony, Warner v. Anthropic (M.D. Tenn., 2023) is an output-infringement case over song-lyric reproduction. RIAA v. Suno and RIAA v. Udio (June 2024) are the principal training-data cases in music: the labels allege that the generative-music services trained on copyrighted masters. Discovery in these cases is producing the first public record of what audio training sets contain.
Code. Doe v. GitHub/OpenAI/Microsoft (N.D. Cal., 2022) on GitHub Copilot is the earliest of the litigation wave, focused on reproduction of copyleft-licensed code without attribution. Most claims were dismissed; a narrowed DMCA § 1202 claim proceeds.
Shareholder derivative actions. A distinct cluster sues corporate officers rather than the company, alleging that directors and officers approved copyright infringement, exposed the company to substantial risk of copyright liability, or made material misrepresentations concealing or minimizing that risk. *Rosen v. Cook*, against Apple chief executive Tim Cook and thirteen other Apple directors and officers over Apple's AI training on copyrighted works, was filed August 14, 2026 in the Northern District of California as No. 5:26-cv-08463, weeks before Cook's retirement at the end of August, and became publicly reported the following day. It is the fifth such US action, following suits against executives of NVIDIA, Microsoft and Adobe on similar theories (Verified Stockholder Derivative Complaint, Rosen v. Cook et al.; Source: chatgptiseatingtheworld.substack.com). The Apple complaint pleads the training-data practices directly — the Books3 pirated-book corpus and the Panda-70M YouTube-derived video corpus — and adds an Illinois Biometric Information Privacy Act count over voice models, where *Anderson v. Microsoft* rested on proxy-statement disclosure. The theory routes around the fair-use question that governs the direct-infringement cases above, resting instead on fiduciary duty and disclosure; no ruling on the merits of any of the five has been reported.
Licensing markets
Three licensing architectures are competing. The first is direct bilateral licensing, exemplified by OpenAI's deals with AP, Axel Springer, FT, News Corp, Reddit, Shutterstock, and others; lab-to-publisher terms are non-public, with reported prices ranging from the low tens of millions to nine-figure multi-year deals. The second is collective rights organizations: the Copyright Clearance Center (US), the CLA (UK), and equivalents have launched AI-training licensing programs that function like blanket rates for institutional users. The third is bottom-up certification and signaling, represented by Creative Commons' 2024 "Preference Signals" work and Fairly Trained's certification regime, which requires licensed or public-domain training data, as an alternative to both lab fair-use claims and collective licensing.
The JCP (Japan Content Platform) and similar national-level licensing bodies are proposed intermediaries for per-country aggregated licensing deals; none had reached operational scale as of 2026.
Settlements and rulings
The 2025 settlement wave narrows what the underlying legal rule has to decide. By April 2026:
- Anthropic and music publishers reached a preliminary settlement on the lyric-reproduction aspects of the UMG/Sony/Warner suit (early 2025); training-data claims proceed separately.
- OpenAI licensing deals with News Corp, Meredith, and Time (late 2024 to 2025) foreclose those publishers as plaintiffs.
- The Getty v. Stability AI UK High Court trial (summer 2025) yielded a split ruling: trademark and database-right claims succeeded, while core copyright claims were narrowed. The ruling is under appeal.
- *Andersen v. Stability* survived Rule 12 and Rule 56 motions and is ongoing through discovery, with 691 docket entries as of July 31, 2026 (Source: courtlistener.com).
- The author claims against OpenAI, led by *Tremblay v. OpenAI* (N.D. Cal., filed June 28, 2023), were transferred by the JPML to the Southern District of New York on April 21, 2025 and consolidated into MDL 1:25-md-03143, the same multidistrict litigation that carries the news-organization claims (Source: courtlistener.com).
None of these has been a decisive fair-use ruling. Whether LLM training on scraped copyrighted data qualifies as fair use remains live in NYT v. OpenAI and Kadrey v. Meta.
A Supreme Court ruling outside the AI docket has been read as bearing on the secondary-liability side of these cases. In March 2026 the Court ruled unanimously in Cox v. Sony, reversing a $1 billion contributory-infringement verdict against the internet service provider and holding that liability attaches only when a company induces infringement or builds a product tailored for it. Electronic Frontier Foundation legal director Corynne McSherry said on July 15, 2026 that the ruling gives AI developers a "clean, clear" defense against claims that models are "infringement machines" in the roughly 100 copyright suits then pending against them (Source: broadbandbreakfast.com).
Separately, Anthropic reached a $1.5 billion settlement with authors and publishers, which addresses the input side of AI training-data copyright (Source: nytimes.com).
Protectability of AI outputs
Separate from ingestion, the US Copyright Office's 2023–2024 guidance and the 2023 Thaler v. Perlmutter ruling establish that purely AI-generated works cannot receive US copyright, because human authorship is required. The 2024 Zarya of the Dawn decision allowed copyright on the human-authored textual elements of a graphic novel but not on the Midjourney-generated images. The result is an asymmetric market in which inputs remain protected but outputs do not. The EU's position is similar but less fully tested.
Enforcement timing and AI-assisted transformation
The speed at which AI systems can reproduce and rewrite protected works has surfaced as a distinct enforcement question. When a Claude Code source-code leak surfaced on GitHub in mid-April 2026, Anthropic invoked copyright and demanded takedown of thousands of original posts. Sigrid Jin then used AI assistants to rewrite the leaked code in a different programming language and re-posted that version; Anthropic did not request takedown of the rewritten version, leaving open whether AI-assisted translation produces a transformative work that escapes the original copyright (Source: nytimes.com). Traditional copyright enforcement assumes a window between unauthorized copying and large-scale propagation; AI agents that can rewrite hundreds of thousands of lines of code, prose, or musical scores in minutes compress that window (Source: nytimes.com). The episode reads against Anthropic's $1.5 billion authors-and-publishers settlement: the settlement addresses the input side of AI training-data copyright (2025), and the leak episode raises the output side (2026).
Relation to policy
Several policy instruments bear directly on the litigation. Training-data disclosure laws, including California AB 2013 and EU AI Act Article 53, reduce the information asymmetry central to litigation. America's AI Action Plan signals a federal preference for a pro-training-fair-use posture, though it is not yet enacted. EU AI Act (Regulation 2024/1689) Article 53 makes training-data transparency a condition of GPAI market access. As a remedy, 17 U.S.C. § 503(b) allows destruction of infringing copies; applied to model weights, this relief is still live in NYT.
The EFF argues the opposite of the plaintiff-side position, and on distributive rather than doctrinal grounds (AI and Copyright: Expanding Copyright Hurts Everyone — Here's What to Do Instead (EFF, February 2025)): new licensing requirements "would lock in the market advantages enjoyed by Big Tech and Big Media — the only companies that own large content libraries or can afford to license enough material to build a deep learning model — profiting entrenched incumbents at the public's expense." Its proposed alternative is instrument-matching: "stronger antitrust rules and enforcement would be a much better" response to the competition concern than copyright expansion.
AI music generation presents the same questions with two differences that have shaped how the cases are pleaded: the training corpus consists of sound recordings whose ownership is concentrated in a few rightsholders, and the outputs compete in the same market as the inputs. Because the developers declined to disclose training data, the RIAA-coordinated complaints establish copying from output evidence instead — including reproduced producer tags, which carry no expressive function and so could only come from the recordings (Complaint, UMG Recordings et al. v. Suno, Inc. (D. Mass., June 24, 2024)).
Separately, complaints have begun pleading web scraping as a reproduction distinct from training, so that a defence to the training question does not dispose of the acquisition question (Class Action Complaint, Hachette Book Group et al. v. Google LLC (S.D.N.Y., July 10, 2026)).
Relationships
- depends-on: AI Copyright Litigation — Analysis — case-level tracking
- supports: NYT v. OpenAI and Microsoft — Complaint (Dec 2023) — the central litigation
- related: EU AI Act (Regulation 2024/1689) — Article 53 GPAI transparency and TDM opt-out layer
- related: Data (Use and Access) Act 2025 — Source Summary — unresolved UK TDM regime
- related: America's AI Action Plan — federal signalling toward pro-training fair use
- related: Synthetic Media / Deepfakes — voice/likeness (ELVIS Act) overlap
- related: Open-Source AI / Open-Weight Models — open-weight release complicates ingestion-liability theories
- related: General-Purpose AI (GPAI) — GPAI is the regulatory category in the EU copyright regime