Web scraping is the automated extraction of content from publicly accessible web services to assemble training corpora. The EDPB's characterization captures why it is regulated as a distinct activity rather than as ordinary data collection: it "is a large-scale processing activity and often occurs without data subjects being aware of it."
Two regimes, one activity
The practice is contested under two bodies of law that ask different questions and can reach different answers about the same conduct.
Data protection asks whether processing personal data was lawful. The EDPB's 2026 guidelines hold that consent is generally unavailable because scrapers "do not have a direct relationship with the data subject," leaving legitimate interest under Article 6(1)(f) as the practical basis, subject to a necessity test that bites directly on scale: "narrowing the collection criteria to exclude unnecessary collection of personal data, rather than scraping a wide part of the internet may be crucial to ensure the necessity condition is met."
Copyright asks whether reproduction was authorized. US complaints increasingly plead scraping as a reproduction separate from training — Hachette v. Google pleads three distinct reproduction counts, one of them specifically for "downloading of web-scraped datasets," stating that this alleges "separate and distinct acts of reproduction" from the training count.
A practice may therefore be lawful under one regime and not the other, and the defences differ: fair use has no analogue in the GDPR, and legitimate-interest balancing has no analogue in copyright.
Consent and robots.txt
The most consequential holding for prevailing industry practice concerns what publication implies. The EDPB states that making data "available online, for example on a web page accessible to everyone, this does not mean that the data subjects gave their consent to the scraping of their personal data for a specific purpose," and — addressing the technical convention the industry has treated as the operative signal — that "the absence or non-applicability of a robots.txt file on a web site does not amount to consent within the meaning of the GDPR."
Platform defaults as a substitute for consent
A distinct configuration arises where the platform holds the content directly and enrolls its creators by default. Twitch said on August 12, 2026 that creators' content will be used to train generative AI models across Amazon, with creators enrolled by default and required to opt out through channel settings. Asked on a company livestream why the setting is not opt-in, chief product officer Mike Minton told an audience of nearly 3,000 users: "If this was opt-in, nobody would opt in. That's honestly the answer." Head of community Mary Kish said the opt-out itself reflects the community's opposition to generative AI training (Source: techcrunch.com).
The stated rationale makes explicit what the EDPB position above addresses in the scraping context: the default is chosen because affirmative consent is not expected to be given. The mechanism differs from scraping in that no crawling is involved and the terms are contractual rather than inferred from publication, which places it closer to the platform-terms route than to the robots.txt question. The pattern of a default-on training or generation setting drawing objection and an opt-out workaround recurs — Meta launched and then discontinued a comparable default-on Instagram-photo setting in July 2026.
What the corpora contain
The composition of scraped corpora is itself an evidentiary question in both regimes. Bender et al. argued in 2021 that scale does not confer representativeness — harassment drives groups off platforms, alternative fora are less linked and so less crawled, and filtering compounds both, since GPT-3's Common Crawl filter selected for documents resembling Reddit-linked pages. The Hachette complaint makes the parallel copyright allegation about the same corpus, citing pirate sources and stating that "the copyright symbol (©) appears more than 200 million times in the C4 dataset."
Downstream safeguards as an input
The EDPB guidelines couple collection to deployment in a way that is unusual for data-protection guidance: mitigating measures relevant to the legitimate-interest balance include "measures to limit the risks of memorisation, regurgitation or attack of AI models or systems." Whether training was lawful can therefore turn partly on how the deployed model behaves — which aligns the data-protection analysis with the output-substitution theories advanced in the copyright cases.
Relationships
- depends-on: EDPB Guidelines 03/2026 on web scraping in the context of generative AI — the principal guidance on when scraping for AI training is lawful under the GDPR
- related: AI Copyright — the parallel regime reaching the same conduct through reproduction
- related: On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? (Bender, Gebru, McMillan-Major, Shmitchell — FAccT 2021), Class Action Complaint, Hachette Book Group et al. v. Google LLC (S.D.N.Y., July 10, 2026), Hachette et al. v. Google (Gemini training data), Training Data Walls, AI and Privacy, EDPB Guidelines 02/2026 on Anonymisation