The GDPR and AI
Be able to apply data minimisation, a legal basis and data subjects' rights to an AI system.
Prerequisites
- EPersonal data and anonymisationrequired
Intuition
The GDPR applies to all processing of personal data — including the training data, the prompts, the logs and the model output.
Six principles that everything else follows from (article 5):
| Principle | What it means in an AI system |
|---|---|
| Lawfulness, fairness, transparency | there is a legal basis and it has been communicated |
| Purpose limitation | data collected for support may not simply be used for training |
| Data minimisation | collect only what is needed — not «it might come in handy» |
| Accuracy | incorrect data has to be correctable |
| Storage limitation | purge according to a stated period |
| Integrity and confidentiality | encryption, access control, logging |
The most common breach in AI projects is purpose limitation: data was collected for one purpose and is then used to train a model. That requires either a new legal basis or that the purpose was stated from the start.
Formal
The legal basis (article 6) — there has to be one, and the choice has consequences:
| Basis | Suits | The catch |
|---|---|---|
| Consent | optional features | has to be as easy to withdraw as to give; invalid where there is an imbalance of power (an employer, a school) |
| Contract | necessary for the service | «necessary» is interpreted narrowly |
| Legitimate interest | much AI development | requires a documented balancing test; the data subject can object |
| Legal obligation | accounting, archiving | |
| Public interest | authorities, schools | most common in the public sector |
For special categories (health, ethnicity, biometrics, political opinion) article 9 applies: processing is prohibited unless an exception applies, in practice most often explicit consent.
Data subjects' rights, and what they mean technically:
| Right | The technical consequence |
|---|---|
| Access (art. 15) | an export function that gathers everything about a person |
| Rectification (art. 16) | the data has to be changeable, including in derived data |
| Erasure (art. 17) | the hardest — see below |
| Data portability (art. 20) | a machine-readable format |
| Objection (art. 21) | under legitimate interest |
| Art. 22 | the right not to be subject to solely automated decisions with significant effects |
Erasure and trained models is the unresolved question. Deleting a person's rows from the database is easy; removing their contribution from a trained model's weights is not. Three practical positions:
- Avoid the problem: do not train on personal data. Pseudonymise or aggregate first.
- Plan for retraining: purge and retrain on a schedule, so that erasure takes effect within a documented time.
- Machine unlearning is an active research area but not a finished solution.
Position 1 is the only robust one, and it should be the starting point.
Three things often forgotten in AI systems:
| Forgotten | Why it matters |
|---|---|
| Prompts are personal data | users type in names, health information, anything at all |
| Logs are personal data | and are often kept the longest of all |
| Third-party providers | a data processing agreement is required, and third-country transfers have to be handled |
A DPIA (a data protection impact assessment, art. 35) is required for high risk — which covers systematic profiling, large-scale processing of sensitive data and monitoring. Most AI systems with personal data end up there.
Interactive
Do a quick GDPR walkthrough of an AI feature. Seven questions; if you cannot answer one of them, that is where the work lies.
| # | Question | Your answer |
|---|---|---|
| 1 | Which personal data is processed? (do not forget the prompts and the logs) | |
| 2 | What is the purpose, stated in one sentence? | |
| 3 | Which legal basis, and why that one? | |
| 4 | What is collected that is not needed for the purpose? | |
| 5 | How long is each category kept, and who purges it? | |
| 6 | How is a request for access or erasure answered? | |
| 7 | Is anything shared with a third party, and is there an agreement? |
Applied to AI-grafen's tutor:
- The pupil's name, email, answers to exercises, dialogue logs with the tutor.
- To provide individually adapted teaching and feedback.
- Contract for the core service; consent for optional parts such as sharing with a mentor.
- Nothing — the dialogue logs are minimised and contain no more than necessary.
- Attempts are kept for the lifetime of the account; the dialogue logs for 90 days; the audit log longer for security reasons.
- An export endpoint that gathers the attempts, the consents and the journal; a deletion flow that removes the account.
- The LLM provider is a data processor — the agreement and the data location have to be in place.
Question 4 is the most valuable and the one that most often reveals something. Almost every system collects something «just in case» that nobody has a purpose for — and every such field is a risk with no matching benefit.
Mastery means
- Chooses a legal basis
- Applies data minimisation and storage limitation
- Handles data subjects' rights in an AI system
Sign in to do the exercises and build your mastery up.
Sources
- GDPR — Regulation (EU) 2016/679 — EU legal act
- IMY — Integritetsskyddsmyndigheten — myndighetsmaterial
- Europeiska dataskyddsstyrelsen (EDPB) — riktlinjer — EU-material