Website scraping: What the new EDSA guidelines mean for companies
Website Scraping: What the New EDSA Guidelines Mean for Businesses
Website scraping – i.e. the automated reading of content from the Internet – is becoming increasingly common, especially through generative AI. But what is data protection allowed? The European Data Protection Authority (EDSA)** has established clear rules for companies and developers with its Guidelines 03/2026. The guidelines distinguish between two types of scraping: **targeted** and **untargeted** scraping.
What is website scraping?
Website scraping refers to the automated process in which a program (so-called *scraper*) extracts data from websites. This can be, for example, the collection of product prices, news articles or contact information. *Generative AI* often uses such data to train responses or create content.
The two types of scraping in focus
The EDSA divides scraping into two categories:
- ** Targeted scraping**: Specific websites or data sources are evaluated here, e.g. to compare prices of competitors. The EDSA sees this as less problematic as long as the data is publicly available and no personal data is affected.
- **Untargeted scraping**: Large amounts of data are collected from many different websites, often without clear selection criteria. This is more critical from a data protection point of view, as personal data (e.g. names, e-mail addresses) can easily be collected here – even if they were not searched directly.
When is scraping data protection compliant?
The EDSA stresses that scraping is only allowed if:
1. **No personal data** will be processed – or the data subjects have given their consent. 2. The data **is publicly available** and no technical protection measures (e.g. *robots.txt* files that prohibit scraping) are circumvented. 3. The purpose of the scraping **transparent** is and the data subjects are informed if their data are affected. 4. The data ** are not stored or processed for indefinite purposes**.
In addition, companies must ensure that they **do not use copyrighted content** without permission. This is especially true if the scraped data is used for the development of AI models.
## What does this mean for your company?
If your company uses website scraping – whether for market analysis, AI training or other purposes – you need to make sure you comply with the EDSA guidelines. In practice, this means:
- Check if the scraped data is **personal**. If yes, you need a legal basis (e.g. consent of those affected). - Avoid bypassing protective measures such as *robots.txt* or Captchas. - Document the **purpose** of the scraping and inform those affected if necessary. - Use scraping tools that are **GDPR compliant** and do not collect unnecessary data.
If you are unsure whether your scraping process meets the requirements, you should seek legal advice. The EDSA guidelines are binding and violations can lead to fines.
## What this means for users
The cloud file component in xynap allows you to store and manage data in a GDPR compliant environment. When you process scraped data in xynap, you can ensure that it is stored locally and under control – without dependence on external providers.