Guidance adopted by the European Data Protection Board on 7 July 2026 under Article 70(1)(e) GDPR, issued for public consultation until 30 October 2026. Primary text: EDPB Guidelines 02/2026 on Anonymisation.
Why the threshold matters
Anonymous data falls outside the GDPR entirely, so the anonymisation test determines the outer boundary of the regulation. For AI specifically it determines how far a developer can rely on anonymisation to remove training-data obligations, and whether model outputs derived from personal data remain in scope.
Key provisions
Anonymity is relative, not absolute. "Whether this is the case may vary from one entity to another. Consequently, anonymity should be assessed from each relevant entity's perspective – typically any party for whom the data is intended to be anonymous." A dataset may therefore be personal data in the hands of the controller who holds the key and anonymous in the hands of a recipient who does not.
Identifiability is tied to differential treatment. A person is identified or identifiable "if they can be distinguished from others in a given context using means reasonably likely to be used and in a way that makes it possible to treat them differently." Naming is not required; the capacity to single out and act differently is.
"Means" is read broadly, and "may include means that are only accessible through a third party," assessed "in light of all objective factors."
Two approaches. The contextual approach "considers the differences in capabilities between those who might identify the data subject" and "reflects the full nuances of the legal standard." The simplified approach ignores those differences; the board states plainly that it "can go beyond the legal standard and may lead an anonymising controller to treat data as though it is not anonymous even if it would actually be so for some relevant entities," offering convenience and confidence in exchange.
The three criteria. Anonymity is tested against No Record Isolation, No Linkage, and No Inference. Re-identification "is more likely to be successful against record-level data with high dimensionality and high resolution." Techniques are judged on whether they "produce an answer which is sufficiently precise and reliable to allow the data subject to be distinguished and treated differently." Passing all three under either approach establishes anonymity; a failure triggers further analysis rather than automatic disqualification.
The guidelines also address mixed datasets, the GDPR status of the anonymisation process itself, and include a glossary and a flowchart for the technical analysis.
Comparison with other approaches
The relative-anonymity holding follows the Court of Justice's reasoning in C-413/23 P EDPS v SRB (September 2025) and departs from an absolute reading under which data is either anonymous or not. For AI training data this matters more than for conventional datasets, because model weights and outputs may be anonymous from a deployer's perspective while the training corpus was not from the developer's.
Read with the companion Guidelines 03/2026, the two texts bracket the lifecycle: 03/2026 governs lawful collection, 02/2026 governs when the resulting data leaves the regime.
Key tensions
High-dimensionality data is where anonymisation is hardest and where machine learning derives most value, so the criteria bite hardest on the datasets most worth building. The simplified approach's admitted over-inclusiveness also creates a compliance asymmetry: an organisation choosing the safer path assumes obligations the law would not impose, while one choosing the contextual approach must defend a per-entity assessment against a regulator.
Relationships
- instance-of: AI Governance (umbrella)
- depends-on: General Data Protection Regulation (GDPR) — interprets the definition of personal data
- related: EDPB Guidelines 03/2026 on web scraping in the context of generative AI, European Data Protection Board (EDPB), AI and Privacy, Three Privacy Problems AI Creates, EDPB Guidelines 02/2026 on Anonymisation