Version 1.0, adopted July 7, 2026 by the European Data Protection Board under Article 70(1)(e) GDPR, for public consultation to October 30, 2026.
Scope
The guidelines "aim to clarify the legal and technical implications of web scraping for generative artificial intelligence (AI) development under the GDPR," covering two scenarios: an organisation that scrapes "themselves or by contracting another party," and one that "obtains and reuses a dataset that has already been scraped by another organisation (data broker)."
Three exclusions bound them. Data brokers who scrape and hold datasets "and whose intention is not to train AI models themselves" are not addressed. Processing of an organisation's own data is out of scope. And because scraping for generative AI "is mainly performed by private sector bodies," the guidelines "focus only on web scraping done by private entities."
The characterization that frames the risk analysis: web scraping "is a large-scale processing activity and often occurs without data subjects being aware of it."
Legal basis
The board's position is that consent is generally unavailable and legitimate interest is what remains, with the other Article 6(1) grounds "generally less likely to apply."
Consent. Organisations scraping third-party data "do not have a direct relationship with the data subject and are most probably not able to identify and obtain consent from each and every data subject before scraping." The board goes further on the inference from publication: "When data subjects make their personal data available online, for example on a web page accessible to everyone, this does not mean that the data subjects gave their consent to the scraping of their personal data for a specific purpose." And the sentence with the most direct bearing on current industry practice: "the absence or non-applicability of a robots.txt file on a web site does not amount to consent within the meaning of the GDPR."
Legitimate interest (Article 6(1)(f)). Three cumulative conditions must be met — a legitimate interest pursued by the controller or a third party; necessity of processing personal data for it; and the balancing test, that the data subject's interests or fundamental rights "do not take precedence."
An interest is legitimate only if "lawful, clearly and precisely articulated and real and present, not speculative." Building on Opinion 28/2024, the board names three examples in the AI context: "developing the service of a conversational agent to assist users"; "developing an AI system to detect fraudulent content or behaviour"; and "improving threat detection in an information system." For a general-purpose model whose use is undecided, controllers are advised to refer to the objective pursued by development — "in particular whether it is commercial, public, scientific research, and whether it is internal or external to the organisation."
Necessity has two elements: whether the processing "will allow the pursuit of the purpose," and "whether there is an equally effective and less intrusive way." Applied to AI training, this bears directly on indiscriminate collection: "narrowing the collection criteria to exclude unnecessary collection of personal data, rather than scraping a wide part of the internet may be crucial to ensure the necessity condition is met," with pseudonymised or synthetic data offered as less intrusive alternatives.
On the balancing test, the data subject's interests include "self-determination and retaining control over their own personal data," and the board identifies a collective harm as well: "large-scale and indiscriminate data collection in the AI development phase may create a sense of surveillance for data subjects ('chilling effect')."
Mitigating measures that can be weighed in the balance include ensuring certain data categories are not collected or certain sources excluded by default, "limiting the collection to freely accessible data," safeguards increasing transparency, facilitating rights exercise, and deleting or anonymising personal data as soon as possible. The board extends the relevant measures downstream: "measures to limit the risks of memorisation, regurgitation or attack of AI models or systems are also relevant" — so training-stage lawfulness can turn on deployment-stage safeguards.
Transparency
The board acknowledges the practical problem — "when scraping large amounts of data from the internet, it is often difficult, impracticable or, even, objectively impossible to identify and inform the data subjects in an effective way" — and then narrows the exemption.
Article 14(5)(b) may excuse individual notice where it "proves impossible or requires disproportionate effort." But the exemption is framed around research: the GDPR names archiving in the public interest and scientific, historical, or statistical research "in particular," and while other cases exist, "this exception should not be routinely relied upon by controllers who are not processing personal data for the purposes of archiving in the public interest, for scientific or historical research purposes or statistical purposes."
Where relied on, the controller "should assess the effort involved… against the impact and effects on the data subject if they were not provided with the information," conducted "with regard to the dataset as a whole and not on every single piece of personal data," weighing the number of data subjects, the age of the data, and any safeguards adopted.
Other principles
Minimisation. Before collection: considering synthetic data instead of personal data; defining precise collection criteria; data mapping and inventory; filters excluding certain data categories; excluding "websites that structurally contain certain types of data and websites which clearly oppose the scraping." After collection: syntax-based filtering and, where feasible, replacement with synthetic data, anonymisation, or pseudonymisation.
Accuracy. Controllers should "scrape from reliable sources, timestamp the data and validate the data before using them in AI training."
Roles. Where several organisations are involved, qualification "as controllers, joint controllers or processors should be analysed on a case-by-case basis."
Special-category data
Processing special categories "is in principle prohibited," requiring an Article 9(2) derogation in addition to an Article 6 basis. The board treats GC & Others (C-136/17) as potentially relevant "for the incidental and residual collection" of such data in AI training, "'within the framework of his responsibilities, powers and capabilities,' if the controller implements technical and organisational measures to prevent the collection and the dissemination of data" — with the qualification that the controller "should carry out a case-by-case analysis of whether the reasoning of the Court's ruling can be applied." This is a narrow accommodation for incidental collection, not a general exemption.
Relationships
- supports: European Data Protection Board (EDPB) — the board's primary guidance on AI training data
- related: EDPB Guidelines 02/2026 on Anonymisation — companion text adopted at the same plenary, governing when anonymisation removes GDPR obligations
- related: EDPB Guidelines 03/2026 on web scraping in the context of generative AI — the legislation page tracking the guidance
- supports: Web Scraping for AI Training — the principal guidance on lawfulness under the GDPR
- related: AI and Privacy, Three Privacy Problems AI Creates, General Data Protection Regulation (GDPR), Training Data Walls