On 8 July 2026, the European Data Protection Board (“EDPB”) issued two related sets of draft guidelines with one being on anonymisation and the other on web scraping for generative artificial intelligence. When read together, they seek to address opposite ends of the same data lifecycle, with the scraping guidelines governing how personal data enters an AI-development pipeline and the anonymisation guidelines explaining when data can credibly leave the GDPR’s scope.

I. Anonymisation

The anonymisation guidelines firmly adopt a relative approach following EDPS v. SRB (Read more on this case by clicking here & here). Data may be personal for one entity yet completely anonymous for another, depending on whether that entity has means “reasonably likely” to be utilised to identify the individual. By way of example, a recipient may receive data which is anonymous from their perspective even though the transferring data controller retains information allowing re-identification of such data, and accordingly must treat its own copy as personal data.

The EDPB asks whether an individual can be distinguished and treated differently through record isolation, through inference or even a link with other information. The proposed framework accordingly sits on three prongs:

  1. No Record Isolation;
  2. No Linkage; and
  3. No Inference.

Passing all three supports a finding of anonymity, albeit failing one does not automatically make the data amount to personal data, but rather, it ought to trigger a deeper assessment of whether the real-world likelihood of identification remains insignificant.

The assessment must inherently account, amongst other things, for the data’s granularity and all available auxiliary information and foreseeable technological advances, the latter being salient with the rapid advancement in AI and its increasing capability to infer.  Interestingly, the capability of the relevant actors is also listed. With regard to such relevant actors, and given that the GDPR does not lay down any conditions regarding the entities who are able to identify the individual, the EDPB provided a non-exhaustive list which includes the likes of: i) rogue employees; ii) investigative journalists; iii) law-enforcement bodies; and iv) malicious third parties.

One thing reads abundantly clear. The state of “anonymity” is not a label to be applied once over data and subsequently forgotten. Circumstances may change and thus, a degree of constant evaluation is expected.

II. Data Scraping

From the outset, the web-scraping guidelines reject any assumption that data published online is available for unrestricted AI training. They apply to private entities scraping external internet sources themselves, outsourcing scraping and even those reusing previously scraped datasets.

Legitimate interests will often be the most realistic Article 6 basis, but the familiar three-stage test remains demanding.[1] (Read more on the three-stage test here). The availability of less intrusive alternatives, such as narrower datasets, synthetic data or effective filtering, directly affects necessity.

Reasonable expectations now acquire a distinctly technical dimension. Robots.txt and ai.txt files,[2] CAPTCHAs, login requirements and express anti-scraping statements may indicate that individuals would not expect their data to be harvested for AI development, albeit their absence should not be construed as tacit consent.

The most difficult issue remains special-category data. Article 9 GDPR looms, and still requires a separate condition. The EDPB leaves limited room for genuinely incidental and residual collection, drawing on GC and Others, but only through a case-specific assessment and lifecycle safeguards (prompt deletion, output filtering, model unlearning).

--

Taken together, the draft guidelines close off two rather convenient assumptions: i) that information published online is therefore available for unrestricted use; and ii) that anonymised data is permanently outside the scope of the GDPR. The regulatory treatment is increasingly employing a pragmatic approach, and any dataset will depend not on what the organisation calls it, but rather what can realistically still be done with it.

Unsure whether your scraping or anonymisation practices meet the EDPB’s emerging standard? For any further information or assistance, please contact us at info@gtg.com.mt

Authors Dr J.J. Galea


[1] Lawfulness, necessity & the balancing test.

[2] Robots.txt is a machine-readable file placed at a website’s root which indicates which parts of the site automated crawlers may or may not access under the Robots Exclusion Protocol.

Ai.txt is an emerging (presently non-standardised) counterpart intended to communicate permissions or restrictions specifically to AI crawlers, including the use of website content for model training

 

Disclaimer This article is not intended to impart legal advice and readers are asked to seek verification of statements made before acting on them.
Skip to content