Personal data and anonymisation
Be able to identify personal data in a dataset and apply pseudonymisation and data minimisation.
Prerequisites
- DLicences and open datarequired
Intuition
Personal data is any information that can be linked to an identifiable natural person — directly or indirectly.
| Type | Examples |
|---|---|
| Directly identifying | name, national ID number, email, phone number |
| Indirect (quasi-identifiers) | postcode + date of birth + sex, IP address, device id |
| Special categories | health, ethnicity, religion, political opinion, sexual orientation, biometrics |
The middle row is the one that surprises people. Sweeney showed that a postcode, a date of birth and a sex are enough to uniquely identify roughly 87 % of the US population. Removing the name is therefore nowhere near enough.
Formal
The distinction that decides what the law requires:
- Pseudonymisation: the identifiers are replaced with an id, but a key still exists (or reidentification is possible). The data is still personal data and the GDPR applies in full.
- Anonymisation: reidentification is impossible even with extra information. The GDPR then does not apply — but the threshold is high, and most «anonymised» datasets are in practice pseudonymised.
The techniques and what they actually give:
| Technique | Protection | Limitation |
|---|---|---|
| Removing the direct identifiers | low | the quasi-identifiers remain |
| k-anonymity | every row is shared by at least k others | vulnerable to homogeneity and background-knowledge attacks |
| l-diversity, t-closeness | variation in the sensitive attributes | complex, destroys utility |
| Differential privacy | a mathematical guarantee via calibrated noise | costs accuracy; ε has to be chosen and justified |
| Aggregation | high if the groups are large | the detail is lost |
Data minimisation is the strongest and cheapest measure: do not collect what you do not need. Data that was never collected cannot leak, cannot be misused and does not have to be deleted.
For AI systems there is the additional fact that the model can memorise training data. Carlini et al. have shown that exact text strings can be extracted from language models. Personal data in the training data is therefore a risk even after the data has been deleted.
Code
import re, hashlib, os
PII = {
"national_id": re.compile(r"\b(19|20)?\d{6}[-+]?\d{4}\b"),
"email": re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+"),
"phone": re.compile(r"\b0\d{1,3}[-\s]?\d{5,8}\b"),
"card_number": re.compile(r"\b(?:\d[ -]?){13,16}\b"),
}
def mask(text: str) -> tuple[str, dict]:
found = {}
for name, m in PII.items():
text, n = m.subn(f"<{name.upper()}>", text)
if n:
found[name] = n
return text, found
PEPPER = os.environ["PSEUDONYM_PEPPER"].encode() # secret, not rotated lightly
def pseudonymise(identifier: str) -> str:
"""A stable id without revealing the original. NOTE: still personal data."""
return hashlib.blake2b(identifier.encode() + PEPPER, digest_size=8).hexdigest()
text = "Kontakta Anna på [email protected] eller 070-1234567."
print(mask(text))
# ('Kontakta Anna på <EMAIL> eller <PHONE>.', {'email': 1, 'phone': 1})
# ↑ the name is not caught by the regex — that needs NER or manual review
The example shows the limitation: regexes catch structured identifiers but not names, places or free text that gives somebody away. Combine them with NER and human review of a sample.
Mastery means
- Identifies direct and indirect personal data
- Applies pseudonymisation and data minimisation
- Understands why anonymisation is harder than it sounds
Sign in to do the exercises and build your mastery up.
Sources
- IMY — Integritetsskyddsmyndigheten — myndighetsmaterial
- GDPR — Regulation (EU) 2016/679 — EU legal act
- arXiv — Extracting Training Data from Large Language Models — arXiv (open access; licence per article)