My thoughts on today’s EDPB stakeholder event on AI models & personal data:
– The EDPB didn’t reveal anything, but the wording of the questions submitted suggests certain assumptions will underlie the upcoming Opinion.
– We had interesting exchanges among stakeholders in Group C, and I heard that there were also lively discussions in the other two groups. From developers of some of the most influential LLMs in the world to academics, with representatives from large and small companies, it was great to have several views represented.
– The recent KNLTB judgment of the EU Court of Justice isn’t yet as widely known as it could be – in particular the implications of its paragraph 51 on the right to object. ( https://lnkd.in/eUi9nJnw )
– The summary?
(i) Personal data in AI models?
Both perspectives were argued (my take: https://lnkd.in/eYXabSfs ).
The synthetic creation of information that appears as personal data was discussed too, as well as the issue of which measures can be taken to limit the risk of being considered as personal data.
There was the question of whether only legal means can be used to test whether a model contains personal data… and surprisingly (in my view, given e.g. the Breyer judgment) there were several attendees suggesting criminals’ perspective should be taken into account.
Various measures were discussed re measures to limit the risk, from pseudonymisation of training data to output-level filtering.
(ii) Balancing exercise & measures to safeguard data subjects’ rights & freedoms re legitimate interests as legal grounds for training?
Reasonable expectations need to be managed, notably through transparency.
�
The focus should not only lie on technical measures but also others, such as the use of synthetic data, pseudonymisation, opt-outs (in our group, we discussed how unrealistic this is at a training phase, though, and that it makes more sense at an output phase), etc.
�
We had a huge first-party vs third-party data discussion, but the EDPB didn�t really cover it in its summary.
�
(iii) Same balancing exercise, but for post-training & fine-tuning phases?
The�difference between developer, deployer and user is important in this case, because the controller may be someone very different – with significant consequences.
Some argued that retraining might be an option, but several in our group emphasised the unrealistic nature of this from a cost and practical perspective.
There were interesting discussions on the exercise of data subject rights, in particular re filtering of output, but proportionality was the key principle discussed in our group in that context, and we discussed the importance of transparency and disclaimers to users.
***
Am I hopeful that this will lead to anything? Unclear, and some clients have already been discussing next steps with me.
We’ll see how the next EDPB stakeholder event goes – I’m attending the “Consent or Pay” one on 18 November as well!
GDPR data protection
Did this analysis get you thinking? Reach out!
DataLaws.net is entirely open-access, and instead of getting your data in exchange for this content, how about another trade? If this commentary saved you research time or sparked an idea, feel free to invite me over for tea, chai or a hot chocolate next time you are around Brussels or Antwerp - or invite me over to your offices for a chat!
Get in touch ↗ Let's connect on LinkedIn ↗