Web Scraping and the GDPR — What the New EDPB Guidelines 03/2026 Say About Generative AI
“If data are publicly available, they can be used freely” — this is one of the most common beliefs about collecting data online. However, from a GDPR perspective, this assumption is incorrect.
On 7 July 2026, the European Data Protection Board adopted Guidelines 03/2026 on web scraping in the context of generative artificial intelligence. The document explains when automated online data collection for AI development and training may comply with the GDPR. It also sets out the conditions organisations using such data should meet.
The Guidelines matter not only to organisations that conduct scraping themselves. They are also relevant to companies that buy or use ready-made scraped datasets, fine-tune models on them, or provide such datasets to others. The key conclusion is simple: public availability does not mean unrestricted use. Likewise, the absence of a technical scraping restriction does not equal the data subject’s consent.
It is also important to remember that Guidelines 03/2026 were adopted as Version 1.0 for public consultation. Their final wording may therefore still change. Nevertheless, they already show how the EDPB interprets the obligations of organisations using web scraping in AI projects.
What Kind of Web Scraping Do the Guidelines Cover?
The EDPB focuses on collecting data from the internet for training and developing generative artificial intelligence.
In practice, the Guidelines cover two main scenarios. The first involves an organisation scraping data itself or outsourcing scraping to another provider. The second concerns using a ready-made dataset previously collected by another organisation.
The Guidelines apply to activities carried out by private entities. They do not focus on data brokers that only provide datasets without using them for model training. Nor do they focus on organisations processing their own data.
The EDPB defines web scraping as automated collection and storage of information from publicly accessible online resources. These may include public registers, open-data portals, news websites, social media, online forums and blogs. The challenge is that such sources very often contain personal data.
Web Scraping May Involve Processing Personal Data
If scraping collects information that directly or indirectly identifies an individual, it involves the processing of personal data.
Therefore, the GDPR may cover not only data collection itself, but also later cleaning, selection, structuring, combining and storage.
Mixed datasets may contain both personal and non-personal information. In such cases, the GDPR applies at least to the part containing personal data.
Scraping for AI purposes creates particular compliance challenges. Organisations may collect vast amounts of information without direct contact with the people concerned. This raises questions about Article 6 legal bases, Article 9 special-category data, Articles 12–14 transparency duties, data minimisation and purpose limitation.
AI models also create additional risks. In some cases, a model may memorise parts of its training data and later reproduce them. It may also allow users to infer new information about specific individuals.
Controller, Processor or Joint Controller?
Performing scraping does not automatically make an organisation a data controller. The EDPB stresses that roles depend on each party’s actual influence over the purposes and means of processing.
For example, an AI developer may outsource scraping and precisely define the sources and data categories to collect. The contractor may then act as a processor, while the developer remains the controller.
The situation differs when an organisation buys or uses a dataset previously collected independently by another entity. In that case, each party may remain responsible for its own separate processing activities.
Joint controllership may also arise when the organisations jointly determine processing purposes and key data collection criteria.
Legal Basis for Web Scraping — Usually Legitimate Interests
One of the key topics in Guidelines 03/2026 is the legal basis for processing data obtained through scraping.
In practice, private organisations will often consider legitimate interests under Article 6(1)(f) GDPR. Other legal bases are often difficult to apply to large-scale online data collection.
Public Data Do Not Mean Consent
With mass scraping, obtaining consent from every person included in a dataset is usually unrealistic. More importantly, publishing information online does not automatically mean consent to its further use for AI model training.
Likewise, the absence of a robots.txt file or other technical blocking measures does not constitute GDPR consent.
Legitimate Interests Require a Full Assessment
Relying on Article 6(1)(f) GDPR cannot be automatic. The controller must carry out a three-step assessment covering:
- the existence of a specific legitimate interest,
- the necessity of processing for that interest,
- balancing the organisation’s interest against the rights, freedoms and interests of individuals.
The Guidelines mention several potentially legitimate interests. These include developing conversational assistants, detecting fraud and improving threat-detection systems.
Balancing Test and Users’ Reasonable Expectations
The balancing test may be the most demanding part of the assessment.
The EDPB emphasises the reasonable expectations of people whose data are collected. Therefore, technical availability alone cannot determine whether scraping is lawful.
Relevant factors may include:
- the nature of the data source,
- how and to what extent the data were made public,
- the platform’s terms and notices,
- the use of
robots.txt,ai.txt, CAPTCHA or other access restrictions, - the relationship between the organisation and the individual,
- the individual’s particular situation, such as age or public-figure status.
Context plays a crucial role here. If a public platform allows automated collection and clearly informs users, this may affect their reasonable expectations.
By contrast, a platform may use technical safeguards and clearly object to AI-related scraping. In such cases, relying on users’ reasonable expectations becomes much harder.
Data Minimisation Also Applies to Large Training Datasets
Large AI models often require extensive training datasets. However, this does not mean that the data minimisation principle stops applying.
The EDPB expects organisations to actively limit the collection of information that is unnecessary for the defined purpose.
Before launching a crawler, organisations should therefore:
- consider using synthetic data instead of personal data,
- define the scope and collection criteria precisely,
- map data categories and sources,
- apply filters that remove unnecessary information,
- exclude sources likely to contain special-category data or content aimed at children,
- consider signals opposing scraping, such as
robots.txt,ai.txtor CAPTCHA.
Minimisation does not end when collection stops. After obtaining the dataset, organisations can apply additional safeguards. These may include regex-based filters for phone numbers or identifiers, synthetic replacements, anonymisation and pseudonymisation.
Transparency — How to Inform People About Scraping
Because web scraping usually collects data indirectly, Article 14 GDPR will generally apply.
Individually notifying millions of people may be impossible or require disproportionate effort. In some situations, controllers may therefore consider the exemption under Article 14(5)(b) GDPR.
However, this does not remove the broader transparency obligation.
The organisation should publish information explaining:
- which categories of data it collects,
- the purposes for which it uses them,
- the legal basis for processing,
- the sources from which it collects the data,
- how the crawler operates.
The EDPB also encourages organisations to describe their sources as precisely as possible. Providing searchable domains or URLs may help. Organisations can also indicate the relevant collection periods.
If an organisation uses a ready-made dataset obtained from another controller, it should also identify that source.
Special Categories of Personal Data — a Particularly Difficult Area
An additional challenge arises when scraped information includes special categories of personal data under Article 9 GDPR.
Processing such data is generally prohibited. Intentional collection therefore requires both an Article 6 legal basis and an Article 9(2) condition.
The difficulty is that organisations cannot always predict which information mass scraping will collect.
Referring, among other things, to GC and Others, C-136/17, the EDPB considers the controller’s actual capabilities. This is particularly relevant to unintended and incidental collection of special-category data.
In practice, organisations should introduce safeguards at different stages:
Before scraping — use appropriate source-selection criteria, filters and exclusions for websites presenting higher risks.
After collection — identify and quickly remove such information, including in response to data subject requests.
During model development — reduce the risk of later extracting personal data from the model and apply output filters.
After deployment — monitor model outputs and, as technologies develop, use techniques that reduce specific data influence.
Simply introducing safeguards is not enough. Under the accountability principle, controllers should also demonstrate that these safeguards work effectively.
What the New Guidelines Mean for Organisations Using AI
For companies developing or fine-tuning generative models, using scraped data or providing such datasets, Guidelines 03/2026 require a structured approach.
In particular, organisations should:
- determine the roles of all parties involved,
- identify and document the legal basis for processing,
- conduct a full balancing test when relying on legitimate interests,
- assess whether the project requires a DPIA,
- provide appropriate information about scraping activities,
- apply source-selection criteria and filters limiting data collection,
- respect technical and other signals opposing scraping,
- provide an effective objection or opt-out mechanism,
- plan how to handle other data subject rights.
GDPR compliance is only one part of assessing whether web scraping is lawful. Copyright, website terms and text-and-data-mining rules may also apply. This includes rights reservation mechanisms under Directive 2019/790.
Common Myths and Mistakes About Web Scraping
“The data are public, so we can use them.” Public availability does not create a legal basis or consent for AI training.
“There is no robots.txt, so scraping is allowed.” Its absence does not amount to GDPR consent or remove the need for legal assessment.
Relying on consent for mass scraping. In practice, obtaining valid consent from every person included in a dataset is usually unrealistic.
No documented balancing test. Simply writing “legitimate interests” in internal documentation is not enough. Organisations must justify the interest, necessity and balancing outcome.
Ignoring objections from website owners or users. robots.txt, CAPTCHA and explicit AI restrictions may influence the assessment of reasonable expectations.
No safeguards for sensitive data. Collecting data without adequate filters may create serious compliance problems under Article 9 GDPR.
Lack of transparency. Organisations should not scrape data without publicly explaining what they collect, where it comes from and why.
GDPR-Compliant Web Scraping — Practical Checklist
- Clarify the roles of all parties involved — controller, processor or joint controllers.
- Identify the appropriate legal basis and document why it applies in the specific case.
- Conduct the full three-step test when relying on legitimate interests — purpose, necessity and balancing.
- Assess individuals’ reasonable expectations, considering the source, publication context and technical restrictions.
- Apply data minimisation when configuring the crawler — define sources, categories, filters and exclusions.
- Consider objections to scraping, including
robots.txt,ai.txt, CAPTCHA and explicit AI-related notices. - Limit collection involving children and special-category data and introduce mechanisms for detection and deletion.
- Publish Article 14 GDPR information, including details about data sources and crawler operation.
- Provide an objection or opt-out mechanism and procedures for handling other data subject rights.
- Assess whether a DPIA is required, especially for large-scale or high-risk processing.
- Document safeguards and test their effectiveness in line with the accountability principle.
- Repeat the assessment periodically, especially when sources, crawler operation or processing purposes change.
Do You Need to Assess the Legality of Web Scraping or Using Data for AI?
Web scraping for generative AI may trigger several GDPR obligations. These range from legal basis and balancing tests to minimisation, sensitive data, transparency and data subject rights.
At Dr Joanna Maniszewska-Ejsmont Law Firm, we support organisations assessing web scraping and dataset use in AI projects. We also conduct Legitimate Interests Assessments (LIAs) and Data Protection Impact Assessments (DPIAs). In addition, we help determine roles and prepare the required GDPR documentation.
If you develop or fine-tune AI models, collect online data or use scraped datasets, explore our GDPR audit and AI compliance services or contact us. We can help assess the risks and structure the process before deployment.

Contact us — we will review your data processing activities from a web scraping perspective.
