AI Policy Wiki
Dashboard

AI in Journalism and Media

medium confidence · updated 2026-06-06

AI's impact on journalism — licensing deals (AP, Reuters, Axel Springer), litigation (NYT v. OpenAI), AI summaries, Perplexity controversies, content detection, and the generative-news risk landscape.

AI relates to journalism in three ways that are in tension. AI developers use journalism text as training data for large language models; AI products act as a distribution intermediary through search summaries and chatbot interfaces; and AI tools serve as production aids for drafting, summarization, translation, and personalization. AI companies train on journalism, then distribute AI-generated summaries that can substitute for journalism traffic, then sell publishers tools to produce more journalism, while the legal framework for compensation is contested. The policy landscape is divided among copyright litigation, licensing deals, platform distribution rules, and public-media guidance.

Scope

The topic spans several distinct uses. Training-data use covers journalism text ingested by LLM developers during pretraining. AI-generated summaries appear in search and chat through Google AI Overviews, Perplexity, ChatGPT, Claude, and others. Licensing and partnership covers agreements between publishers and AI companies. Litigation covers copyright infringement and unfair competition claims. Newsroom production tools cover AI for drafting, translation, and transcription. Synthetic-news risks cover AI-generated fake stories, images, and video, including deepfakes. Detection and authentication covers C2PA provenance and AI-content labeling.

Copyright analysis treats two of these uses as distinct. Training-time use implicates whether ingestion of copyrighted text is fair use. Inference-time use, such as retrieval-augmented generation that cites publisher content, raises separate issues around attribution, substitutability, and secondary liability.

Licensing and litigation

Publishers have split roughly into two postures toward AI companies. Some have licensed their content; others have sued.

Publishers that have signed licensing deals, typically with OpenAI and increasingly also with Anthropic and Google, include the Associated Press, Axel Springer, Reuters, News Corp, the Financial Times, Condé Nast, The Atlantic, Vox Media, and Time. Publishers that have sued include The New York Times, the most prominent plaintiff, along with the Intercept, Raw Story, the Center for Investigative Reporting, and the Daily News. The two postures reflect different strategic calculations: licensing first-movers receive payment and referral traffic, while litigants seek precedent-setting damages. Outcomes in pending cases are expected to shape the equilibrium.

Disclosed and reported terms of the major licensing deals vary by archive quality, brand, and ongoing content volume. The Associated Press–OpenAI agreement (July 2023) was the first major publisher-AI licensing deal, giving OpenAI access to AP archive text and AP access to OpenAI technology. The Axel Springer–OpenAI agreement (December 2023) licensed Politico, Business Insider, Bild, and Welt content to ChatGPT, with attribution and linking in responses, reportedly worth more than $10 million a year. The Reuters–Meta agreement (October 2024) was a real-time news partnership for Meta AI. The News Corp–OpenAI agreement (May 2024) was a five-year deal reportedly worth $250 million covering the Wall Street Journal, the New York Post, The Times (UK), and others. The Condé Nast–OpenAI agreement (August 2024) covered the New Yorker, Vogue, Wired, and others. Aggregate publisher AI-licensing revenue, still opaque, appears to reach hundreds of millions of dollars annually across the industry.

New York Times v. OpenAI and Microsoft

The New York Times complaint, filed in December 2023, is described as the central AI-copyright case for journalism. The Times alleges that OpenAI scraped millions of NYT articles for training; that GPT models can reproduce substantial verbatim excerpts, with the complaint showing dozens of pages of near-verbatim reproduction; that the outputs substitute for original NYT content, constituting unfair competition; and that Microsoft is secondarily liable as the Bing Chat operator and an OpenAI investor. Relief sought includes destruction of models trained on NYT content. OpenAI's defense relies heavily on fair use. The case is scheduled for bench-level rulings through 2026 and is expected to set the doctrinal baseline for AI-copyright analysis. See AI Copyright Litigation — Analysis for broader context.

Perplexity disputes

Perplexity has been the focus of sustained journalism controversy. In June 2024 Forbes accused Perplexity of near-verbatim reproduction of a Forbes investigative story in a Perplexity "Page" with minimal attribution; Perplexity apologized and agreed to a revenue-sharing program. Also in June 2024, Wired, owned by Condé Nast, reported that Perplexity was scraping content despite robots.txt exclusion, including from sites that had explicitly blocked Perplexity's crawlers. In October 2024 Dow Jones, publisher of the Wall Street Journal, and the New York Post sued Perplexity, alleging systematic copying, and the New York Times sent Perplexity a cease-and-desist demanding it stop using NYT content. Perplexity's subsequent Publisher Program offers a revenue share for publishers whose content appears in answers, an example of a company converting litigation pressure into a licensing business model.

Distribution and traffic

A central empirical question is whether AI-generated summaries, including Google AI Overviews, Perplexity answers, and chatbot responses that synthesize news, substitute for publisher traffic or refer to it. Early evidence is mixed. Publishers report referral-traffic declines from AI Overviews, while search platforms argue that summaries increase user satisfaction and sustain overall search use. The answer bears on whether licensing deals adequately compensate publishers for the distribution change. Some publishers with deals report that the calculus remains negative when traffic lost to AI summaries is taken into account.

Newsroom production tools

Publishers have themselves deployed generative AI for content production, with results that vary by content type. Bloomberg produces AI summaries of earnings reports with human review. The Washington Post operates Climate Answers, an AI chatbot for climate reporting, among other AI-assisted tools. The Associated Press has used AI for earnings reports since 2014 and for sports recaps. CNET retracted multiple AI-drafted articles in 2023 after they were found to contain errors and plagiarism. Sports Illustrated published content under fictitious AI-generated bylines in 2023, triggering newsroom fallout. The pattern across these cases is that AI production tooling has performed better on structured, data-driven content than on flagship editorial work, where failures have been visible.

Detection and authentication

The content-detection tools used in education (see AI in Education) — GPTZero, Originality, and Turnitin — have also been deployed by journalism organizations and platforms to identify AI-generated text. Accuracy is low, false-positive rates are high, and defensibility in disputes is limited. Provenance approaches such as the C2PA content-authentication standard remain industry-voluntary.

Public-media policies

The BBC has been among the most cautious major broadcasters. In 2023 it blocked OpenAI's GPTBot crawler and published principles for AI use requiring editorial oversight and transparency and prohibiting AI for generating factual news without human verification. NPR, CBC, and other public broadcasters have adopted similar postures. The public-media sector has set more cautious norms than commercial publishers.

Synthetic-news risks

Distinct from accidental error is the deliberate use of generative AI to produce fake news at scale. NewsGuard has catalogued hundreds of AI-generated fake-news websites, described as "pink slime" operations, producing content at scale. Russian, Chinese, and Iranian influence operations use generative AI to produce and translate content at volume. The interaction between platform-algorithm amplification and the supply of generative content is not yet well characterized. See AI in Elections and Democratic Institutions for the election-integrity overlap.

Policy responses

The principal policy instruments and their status are:

InstrumentScopeStatus
Copyright litigationAI training useActive; dispositive rulings expected 2026
Licensing dealsArchive + real-time newsExpanding; opaque terms
Platform rulesAI content labelingUneven
EU Copyright DirectiveText-and-data-mining opt-outIn force; AI Act reinforces
EU AI Act transparencyTraining-data disclosureIn force Aug 2025 (GPAI provisions)
C2PA / provenance standardsContent authenticationIndustry-voluntary
Public-media policiesEditorial use of AIInstitution-by-institution

The EU approach differs structurally from the US. The EU Copyright in the Digital Single Market Directive (2019) creates a text-and-data-mining right with an opt-out mechanism that publishers can invoke. The AI Act reinforces this by requiring GPAI providers to respect EU copyright opt-outs and publish training-data summaries; its GPAI provisions came into force in August 2025. US law has no equivalent structure, and the question is litigated case by case.

Relationships