The Irish Data Protection Commission’s statement of 21 May 2025 “on Meta AI” (just like the judgment coming today in a German court case) provides a great excuse to talk briefly about GDPR legal grounds & AI model training, while I continue to work on the next part of my Better Regulation series (that next instalment will notably look specifically at Article 6 GDPR and whether any reform is needed to that provision).
Readers may already be familiar with my piece of last October, “AI training data = (non-)personal data? And is consent really relevant?”. Since its publication, the European Data Protection Board (EDPB) has published its own Opinion on personal data & AI models[1] and there has been more recently a flurry of activity by supervisory authorities keen to tell everyone in the EU how to object to Meta’s intention to train its AI model on content shared by users with the world.
This seems like a great context in which to look again at whether – assuming there is personal data being processed – there are legal grounds justifying the training of AI models on publicly available information (and again with the important question of whether first-party data should be treated any differently than third-party data).
In This Article
- I. A word of caution regarding market interference
- II. AI training & public data
- A. Is public data such as Common Crawl data “personal data”?
- B. Pre-training measures
- C. Other perspectives, and consequences
- III. Legitimate interest assessment – and whether first-party vs third-party should have any influence
- A. Legitimacy test
- B. Necessity test
- C. Balancing test – and what about historic data?
- D. Conclusion: legitimate interests can be a valid legal ground
- IV. The dangers of assumptions by campaigners, regulators and courts
I. A word of caution regarding market interference
A first thing I want to get out of the way is the awkwardness of supervisory authorities telling data subjects what to do or not to do in relation to a specific organisation.
The recent statements in April and May 2025 by supervisory authorities across Europe (see for example Belgium [French / Dutch], France, Germany [such as Baden-Württemberg, Hamburg], Italy, Netherlands, Norway, Poland, Portugal, Sweden) seem unheard of.
While none of them explicitly say “use your right to object” but rather “if you wish to object, here is how to do it”, the fact that this pan-European flurry of statements has happened, while it did not happen for instance in relation to certain other platforms when they rolled out their AI model training (and the statements do not call out any other platforms or means to object there, either), suggests to data subjects that this is something that they should do in this particular case.
It is not the first time that authorities call out individual companies in positive or negative terms, though, as the CNIL did when it was specifically endorsing certain analytics providers as an alternative to Google Analytics. [I recall their list initially only including three or so providers; now there appear to be 23 on the list.]
The flurry of statements when the AI model DeepSeek and its related app were launched also suggested a desire to be seen to say something, though there was no comparable exhortation to data subjects to exercise certain rights.
I do believe that supervisory authorities should err on the side of caution, though. These kinds of publications that talk about individual organisations favourably or unfavourably constitute a form of market interference by public authorities. Talking about a decision seems like the sort of publication that could be justified in certain cases, but talking about a business practice that has not even been investigated by that authority? This seems like having prejudged a situation. [The Dutch authority for instance said literally “It is the question whether this is permitted” in its statement.]
Now the Irish DPC, the authority that has apparently had a dialogue with Meta on the topic, has come out with a public statement that suggests that many or even all concerns have already been looked at. The DPC stresses that the training would only concern “public content shared by adults” on Facebook and Instagram and that there are “measures to protect data subjects, such as de-identification, filtering of data sets and output filters”. The DPC doesn’t exhort users to object, and it doesn’t even include links to the means to object.
Which raises interesting questions regarding public data and how AI training is taking place today.
II. AI training & public data
The typical body for training AI systems constitutes information published on the web. Common Crawl is a frequently used source, a “free, open repository of web crawl data” maintained by a non-profit.
In terms of phases, information is collected from various sources (such as Common Crawl), is prepared (a form of “pre-processing”) for training and is fed to the AI model for training. Afterwards, a user interacts with the AI model (directly through a user interface or indirectly, through a separate application using an API/application programming interface) to generate output. These two broad phases – training and then use – are important to bear in mind, as are various perspectives: the perspective of the publisher making information publicly available on behalf of the data subject, the perspective of the AI model provider and the perspective of the user of the AI model.
A. Is public data such as Common Crawl data “personal data”?
As argued in my previous analysis, there are reasons to consider that the broader training phase (getting a copy of such a repository of information, preparing it for AI training and then the actual training) does not constitute processing of personal data by the AI model provider and that even the use of the AI model by the user does not involve the processing of personal data from the training data from the perspective of the AI model provider.
Where information contained in public data actually relates to a natural person, it is not typically an identified natural person (save perhaps for raw [= not yet filtered] first-party data – more on that hereunder), so the fallback under the GDPR become the identifiable natural person criterion. That criterion has to be understood in the light of Recital 26 of the GDPR as well as the Breyer[2] and Scania[3] judgments of the EU Court of Justice, which in practice mean that there must be lawful and proportionate means at the disposal of the AI model provider to directly identify the natural person or to compel someone else to identify the natural person (= indirect identification). The SRB case might lead the EU Court of Justice to clarify this even further.
In practice, no AI model provider using third-party public data is normally in a position to do this, and even AI model providers using first-party public data will not be in a position to do this.
B. Pre-training measures
This is because of the need to prepare the dataset for ingestion by an AI model.
This leads me to an important clarification: the objective of AI model training is not to provide responses to users that contain personal data but rather to help the AI model predict what a response through the user-facing AI system should look like. [I am not talking here about the use of Internet search capabilities that are being built on top of certain AI models, where the intention is also to fetch information and then include that information in the response, alongside the generated parts of the response. That merits a separate assessment.]
It is precisely for that reason that AI model providers deploy a range of techniques on training data, techniques such as the “de-identification, filtering of data sets and output filters” to which the DPC alludes. De-identification and data set filtering allow the content going in to be less likely to contain irrelevant data, while output filters limit the likelihood of the output, i.e. content going out, being irrelevant.
The interesting thing here is that these measures (i) limit the possibility to consider training data as personal data and simultaneously (ii) limit the risk associated with processing of personal data if one were to consider that it is processing.
C. Other perspectives, and consequences
I have not yet been convinced by counterarguments to the point I make about public data not being personal data from the perspective of the AI model provider.
The typical counterarguments include “surely it must be personal data for BigTech” or “personal data must be interpreted broadly”, but the combination of Recital 26 GDPR + Breyer + Scania does not in fact lead to the conclusion that training data is personal data for the AI model provider just because the data could be personal data for someone else like the publisher or the data subject.
The counterexamples used are always based on public figures whose information appears repeatedly in the training data (or are examples that perfectly illustrate that information will be added to match for instance the pattern of a birthdate, even if the date is incorrect), but those alleging it is personal data do not show why the information being processed isn’t actually synthetic data that I as a user (and not the AI model provider) interpret based on my own context as being personal data.
The EDPB’s Opinion on AI models is very light on details in this respect, and it even mixes up the AI model provider’s perspective both in the training phase and in the use phase with the “perception of the output by the user” part.
And this gets to the issue of what to do with public data. If data has been published online and remains online:
- Why should the data be accessible to everyone for consultation, but not for AI model training?
- Is there really a difference to be made between data based on the location on which it is published if it is public?
The first question is tied to the issues of reasonable expectations and purpose limitation, points that I will get back to hereunder.
The second question is related to the issue of “first-party data” versus “third-party data”, which merits consideration.
In the example of any social media network, if information is made public, anyone can view it. I am not talking about situations where someone makes content only available to their connections/followers/friends/… but where the information is public. If information is made public, even people without an account might be able to access the content depending on how a platform works.
I have seen situations where social media content is used, not only by the platform providers themselves but also by third parties, just like any other scraped data.
So it is worth examining whether this public data can at all be used for AI model training by anyone – not just by the social media platform provider but also by anyone else.
III. Legitimate interest assessment – and whether first-party vs third-party should have any influence
Which legal grounds enable the use of public data for AI model training? The answer is fairly obvious: legitimate interest is the default legal ground for the use of public data for AI model training. Consent doesn’t make sense for public data, and even for first-party data it hardly makes sense, as requiring consent in the case of first-party data would in effect prevent the platform provider from relying on legitimate interests while a third party would be free to invoke that legal ground.
[As mentioned in the previous analysis, some courts have suggested otherwise, for instance a German court in Bayreuth in 2018, which held that legitimate interests could not be relied upon as a legal ground if consent was an option. This came from its assessment of the “necessity” test, during which the court stated that “[i]t must therefore also be taken into account that the applicant acquires the data in particular in the context of ordering processes and that it would therefore be possible for it to obtain consent to the transmission of the data to [a given recipient] in individual cases without disproportionate effort”[4] (rough translation). I’ll get back to this point in relation to the KNLTB case hereunder.]
In its AI models Opinion, the EDPB included the outline for a legitimate interest assessment in relation to AI model training.
It concluded the following:
- There is a legitimate interest;
- Necessity of the processing requires demonstration;
- There can be a balance in favour of the controller, but that requires measures to safeguards data subject rights.
There are many nuances to this position, in particular as regards the necessity and balancing tests, and it is worth going through the EDPB’s positions and possible interpretations or challenges.
A. Legitimacy test
Yes, there is a legitimate interest according to the EDPB:
“the following examples may constitute a legitimate interest in the context of AI models: (i) developing the service of a conversational agent to assist users; (ii) developing an AI system to detect fraudulent content or behaviour; and (iii) improving threat detection in an information system” (para. 69 of the EDPB Opinion)
The biggest issue with this list is that AI models are not the end-result but rather the intermediary step towards the end-result. It is the application of the AI model that provides the purpose. An AI model used within a chatbot meets the first interest highlighted by the EDPB. But in and of itself, the AI model does not have a singular purpose, because it is not usable in and of itself. This reinforces again the question as to whether there is even processing of personal data in the context of training of an AI model, as there is no “purpose” as such – the result or integration defines the true purpose.
B. Necessity test
On the necessity test, there could be necessity, according to the EDPB, but only if one can show that the pursuit of the purpose is not possible “through an AI model that does not entail processing of personal data” (para. 73 of the EDPB Opinion).
Some have interpreted this as meaning that only synthetic data can be used to train AI models – but many leave out the part that synthetic data requires actual data first in order to be created. See therefore my earlier point about whether there is even any processing of personal data – and ways to limit the risk of processing in any event.
The EDPB stated here as well that “[t]he existence of means that are less intrusive to the fundamental rights and freedoms of the data subjects may vary depending on whether the controller has a direct relationship with the data subjects (first-party data) or not (third-party data)” (para. 73 of the EDPB Opinion), with a reference to the strange section in the KNLTB judgment in which the EU Court of Justice seemed to suggest that for legitimate interests to be invoked by a first-party controller, that controller might have to “ask [data subjects] whether they want their data to be [processed] for [the relevant] purposes”[5].
That KNLTB passage is extremely odd and even at odds with the very idea of legitimate interest as a legal ground. It transforms legitimate interest as a legal ground into something that requires a consent-type question, which brings into question its very feasibility as a legal ground for a number of processing activities (such as dispute management, service improvement, etc.). After all, if a controller needs to ask a data subject if they wish their data to be processed for given purposes, where is the difference between consent and legitimate interests? Is the only difference the fact that the default option can be “yes, that’s fine”?
While on paper this may sound like a good idea, it means that it becomes easier and more interesting to use third-party data than first-party data – whether in the context of training an AI model or of looking into employee performance, dispute management or customer relationship management. This is not workable and cannot have been the intention of the legislator.
For this reason, I believe that the EDPB’s own interpretation of the necessity test should not be read too literally (and I believe that this part of the KNLTB judgment will be revisited by the CJEU in future cases). It should not mean that if a synthetic dataset is available, it is the only option. Instead, the AI model provider should be able to prove that it has taken measures to limit as far as is (commercially) reasonable the risk associated with the dataset, so as to limit its intrusiveness.
C. Balancing test – and what about historic data?
In its balancing test exercise, the EDPB highlighted various fundamental rights and freedoms of data subjects and how they might be impacted by the deployment of an AI model (paras. 77-90 of the EDPB Opinion).
It then looked at the reasonable expectations of data subjects.
Campaigners today claim that there are no reasonable expectations of data subjects regarding “historic data”, but the EDPB did state in its Opinion that “it is important to consider the wider context of the processing. This may include, although is not limited to, whether or not the personal data was publicly available, the nature of the relationship between the data subject and the controller (and whether a link exists between the two), the nature of the service, the context in which the personal data was collected, the source from which the data was collected (e.g. the website or service where the personal data was collected and the privacy settings they offer),the potential further uses of the model, and whether data subjects are actually aware that their personal data is online at all” (para. 93 of the EDPB Opinion).
While information regarding the processing of personal data has to be provided before the processing occurs, nothing in the GDPR requires “reasonable expectations” to be present before a data point exists. The reasonable expectations have to exist before the processing (or more precisely: before the collection for that purpose) occurs, not before the data exists. Otherwise, legitimate interest could never be relied upon for any new processing of personal data. Recital 47 GDPR emphasises the need to examine “whether a data subject can reasonably expect at the time and in the context of the collection of the personal data that processing for that purpose” – and the collection for that purpose occurs today, not in the past (even in the case of third-party data). [Even the EDPB’s newest guidance on legitimate interests[6] does not suggest otherwise.]
In fact, objection rights exist precisely for that reason. They are there to prevent the processing of both new and old data for a given purpose.
And if we look at the types of measures that the Irish DPC has looked at regarding Meta specifically, and more generally at the list of mitigating measures that the EDPB lists in its Opinion, we can see a list of the types of measures that would help all AI model providers to ensure that a balancing exercise takes data subject rights and freedoms into account. The measures mentioned by the EDPB notably include:
- pseudonymisation (= the de-identification measures to which the DPC referred);
- the replacement of personal data with fake data;
- leaving a “reasonable period of time between the collection of a training dataset and its use”, notably to allow data subjects to exercise their rights;
- proposing an “opt-out” possibility;
- allowing notices to the AI model provider in case of personal data regurgitation or memorisation;
- transparency measures, such as media campaigns and information going beyond the strict requirements of Articles 13 & 14 GDPR;
- measures to facilitate or accelerate the exercise of individuals’ rights in the deployment phase.
Each of these measures contributes to ensuring that data subjects’ rights and freedoms are taken into account. Any useful combination should thus also be taken into account by supervisory authorities, courts and potential claimants.
D. Conclusion: legitimate interests can be a valid legal ground
It is clear then that, based on the EDPB’s own Opinion and by avoiding interpreting it in a way that makes any AI model training whatsoever unworkable, legitimate interests can be used as a valid legal ground for AI model training under the GDPR.
IV. The dangers of assumptions by campaigners, regulators and courts
One element that I now think is worth emphasising is the risk to legal certainty and businesses that is being presented by some recent evolutions.
The market interference risk I highlighted above is not the only thing going on. Campaigners are launching claims left and right (some out of frustration with regulators), with now cease-and-desist actions becoming ever more frequent, some seemingly based on rumour and assumptions.
Among the claims I have seen being alleged recently, some campaigners have been claiming that AI model training by certain organisations has to be stopped because of:
- an alleged lack of reasonable expectations because this concerns “historic” data (despite many organisations being transparent about this and allowing simple tools to object in advance, and despite the EDPB not even suggesting that this is an issue),
- unsubstantiated fears of processing of special categories of data (this isn’t entirely new as a claim – I have seen similar unfounded fears regarding online advertising for several years),
- the allegation that data subjects should be able to review the legitimate interest assessment (if regulators or courts give in to such a claim, they are empowering data subjects to know everything about a business when no other area of law has ever granted such sweeping intrusions into corporate documentation) and
- the claim that this is in breach of the purpose limitation principle and of the compatibility test (though many do not base this on the compatibility of purposes but instead perform a new collection, so an entirely new assessment).
Based on all that I have seen at a large number of organisations that are active in the space, though, these claims tend to be incorrect.
Unfortunately, many are quick to repeat these claims without verifying them. In addition, it may happen that some courts and regulators ignore technical and legal counterarguments because the claims are simple while the truth is more technical, more complex, nuanced, and those flawed decisions then become precedents or authoritative for the future.
Yet the reality is that businesses are looking to integrate data protection concerns into their processes. AI model training is made as privacy-friendly as can be, through safeguards that are put in place at various stages.
It is therefore critical for all in this space, from campaigners to judges, and every data protection professional in between, to carefully consider each assumption or claim being made – assumptions and claims by the businesses themselves, of course, but also the assumptions and claims made by complainants. Accountability only starts when someone is actually a controller (and thus that there is actually processing of personal data), and accountability does not mean a right for data subjects to get everything.
There are limits to the right to data protection, after all, and while not everything is permitted with AI, only the things that are actually unlawful are prohibited.
This was an intermezzo in the context of my new Better Regulation series, which examines ways in which rules can evolve and be improved. For more information on that series, read:
- My take on rethinking the ePrivacy Directive
- My suggested improvements to Articles 1, 2, 4(1), 5 and 30 of the GDPR
[1] European Data Protection Board (EDPB), Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models, 17 December 2024.
[2] Court of Justice of the European Union (CJEU), Breyer, 19 October 2016, C‑582/14, EU:C:2016:779.
[3] CJEU, Scania, 9 November 2023, C-319/22, EU:C:2023:837.
[4] VG Bayreuth, Beschluss vom 08.05.2018 – B 1 S 18.105, para. 72. Original in German: “Zu berücksichtigen ist daher auch, dass die Antragstellerin die Daten insbesondere im Rahmen von Bestellvorgängen erwirbt und es ihr deswegen ohne einen unverhältnismäßig großen Aufwand möglich wäre, im Einzelfall eine Einwilligung zur Übermittlung der Daten an F. einzuholen.”
[5] CJEU, Koninklijke Nederlandse Lawn Tennisbond (KNLTB), C‑621/22, EU:C:2024:858, para. 51.
[6] EDPB, Guidelines 1/2024 on processing of personal data based on Article 6(1)(f) GDPR, 8 October 2024, see page 16.
Did this analysis get you thinking? Reach out!
DataLaws.net is entirely open-access, and instead of getting your data in exchange for this content, how about another trade? If this commentary saved you research time or sparked an idea, feel free to invite me over for tea, chai or a hot chocolate next time you are around Brussels or Antwerp - or invite me over to your offices for a chat!
Get in touch ↗ Let's connect on LinkedIn ↗