Guidance adopted by the European Data Protection Board on 7 July 2026 under Article 70(1)(e) GDPR, issued for public consultation until 30 October 2026. Primary text: EDPB Guidelines 03/2026 on web scraping in the context of generative AI.
Status and instrument type
The guidelines are EDPB interpretive guidance rather than binding legislation. They bind indirectly: national data protection authorities apply the GDPR, and EDPB guidelines are the reference against which supervisory authorities and the European Data Protection Supervisor assess controller conduct. Version 1.0 is a consultation draft, so provisions may change before final adoption.
Scope
The guidelines cover two configurations: an organisation that scrapes "themselves or by contracting another party," and one that "obtains and reuses a dataset that has already been scraped by another organisation (data broker)." Three exclusions bound them — pure data brokers who do not train models, processing of an organisation's own data, and public-sector scraping, since the practice "is mainly performed by private sector bodies."
Key provisions
Legal basis. Consent is treated as generally unavailable, because scrapers "do not have a direct relationship with the data subject." Legitimate interest under Article 6(1)(f) is the practical basis, subject to three cumulative conditions — a legitimate interest, necessity, and the balancing test. An interest qualifies only if "lawful, clearly and precisely articulated and real and present, not speculative."
Publication is not consent. Two holdings bear directly on prevailing industry practice: making data "available online, for example on a web page accessible to everyone, this does not mean that the data subjects gave their consent to the scraping of their personal data for a specific purpose"; and "the absence or non-applicability of a robots.txt file on a web site does not amount to consent within the meaning of the GDPR."
Necessity constrains indiscriminate collection. "Narrowing the collection criteria to exclude unnecessary collection of personal data, rather than scraping a wide part of the internet may be crucial to ensure the necessity condition is met," with pseudonymised or synthetic data named as less intrusive alternatives.
Deployment-stage safeguards enter the training-stage balance. Mitigating measures relevant to the balancing test extend downstream to "measures to limit the risks of memorisation, regurgitation or attack of AI models or systems" — so whether training was lawful can turn on how the deployed model behaves.
Transparency. The Article 14(5)(b) exemption from individual notice, where provision "proves impossible or requires disproportionate effort," is framed around archiving and research purposes and "should not be routinely relied upon by controllers" outside them. Where relied on, the balancing "should be carried out with regard to the dataset as a whole," weighing the number of data subjects, the age of the data, and safeguards adopted.
Special-category data. Article 9 processing remains prohibited absent a derogation. The board treats GC & Others (C-136/17) as potentially relevant to "incidental and residual collection" where the controller implements measures to prevent collection and dissemination — a narrow accommodation, not a general exemption, and subject to case-by-case analysis.
Comparison with other approaches
Where the EU AI Act regulates AI systems by risk tier and imposes training-data governance duties on general-purpose model providers, these guidelines reach the same activity through data-protection law and attach obligations to the controller irrespective of the model's risk classification. They are the EU counterpart to the copyright-based challenges to scraping pursued in US litigation (Hachette et al. v. Google (Gemini training data), UMG Recordings v. Suno (AI music training data)), which contest the same conduct on reproduction rather than personal-data grounds — meaning a scraping practice may be lawful under one regime and not the other.
The companion Guidelines 02/2026 on Anonymisation, adopted at the same plenary, set the threshold at which GDPR obligations cease, and so determine how far anonymisation can serve as the mitigation these guidelines contemplate.
Key tensions
The necessity limb and the scaling logic of frontier training pull against each other: model developers have argued that breadth of data is what produces capability, while the guidelines treat breadth as presumptively unnecessary absent justification. The transparency provision creates a similar bind — individual notice is impracticable at scale, but the exemption designed for impracticability is expressly not available as a routine matter to commercial developers.
Relationships
- instance-of: AI Governance (umbrella)
- depends-on: General Data Protection Regulation (GDPR) — interprets Articles 5, 6, 9, and 14
- related: EDPB Guidelines 02/2026 on Anonymisation, EU AI Act (Regulation 2024/1689), European Data Protection Board (EDPB), Web Scraping for AI Training, AI and Privacy, EDPB Guidelines 03/2026 on web scraping in the context of generative AI