AI models and the notion of “personal data”: a relativistic approach

While I disagree with his conclusions, do read the post of Damien Desfontaines re the BfDI’s public consultation on personal data in AI models. Pity that it ends on 31 August, with the EU Court of Justice’s SRB judgment coming out on 4 September.

I hope the BfDI takes a relativistic approach to the notion of “personal data” into account – especially if the CJEU confirms that relative view (which in my view it already did through its Breyer, Scania and IAB Europe judgments).

Why then link below to Damien’s analysis? Because a proper debate is needed. Not with soundbites but with an in-depth legal discussion.

Here is a summary of arguments I previously set out in favour of an alternative approach:

1. “Personal data” depends on whether a person is (a) identified by the controller [= this concerns that guy, Peter] or (b) identifiable (i.e. that the controller cannot identify directly but indirectly, through data combination or by requesting another to enable identification).

2. This “identifiable” part is where Recital 26 GDPR and the CJEU’s Breyer judgment come into play: someone is identifiable only if the controller has *lawful and reasonable* means at its disposal to get identification (a) on its own or (b) through someone else.

3. When using data collected through web scraping (directly or through e.g. Common Crawl), identification of a person is *NOT* reasonable. No matter an AI provider’s size, none are able to analyse every webpage coming their way and determine (a) whether the contents relate to a natural person and (b) whether that person is identifiable on the basis of the information on that page or in combination with other information.

4. Memorisation is a red herring in this context, because it does not look at identifiability (= key criterion for determining if something is personal data) but at the likelihood of replication of information *based on the user’s request for something that will be viewed by the user as personal data*. Same with the risk of reidentification by threat actors – cyber attacks are not *lawful* means!

5. From the AI model’s perspective everything is created – i.e. *synthetic data*, not personal data. Anecdotal examples (“Donald Trump’s birthday is always correct, so it’s personal data!” or “Max Schrems’s date of birth is incorrect, so that’s inaccurate personal data!”) always look at the perspective of the user, not that of the AI model provider.
Just because the *user* thinks something looks like a duck & quacks like a duck doesn’t mean it’s one for the provider.

6. “But scale cannot help avoid the GDPR”: the focus here is the analysis of (e.g. linguistic) patterns and probabilities, not the processing of discrete data items. Context matters, especially for assessing what is “personal data”.

It’s an important discussion to have.

For more on this, read my previous articles (https://lnkd.in/eQRFZxCi and https://lnkd.in/eYXabSfs).
Damien’s post: https://lnkd.in/eC-HH4-V

Data protection privacy

🫖

Did this analysis get you thinking? Reach out!

DataLaws.net is entirely open-access, and instead of getting your data in exchange for this content, how about another trade? If this commentary saved you research time or sparked an idea, feel free to invite me over for tea, chai or a hot chocolate next time you are around Brussels or Antwerp - or invite me over to your offices for a chat!

Get in touch ↗   Let's connect on LinkedIn ↗