Skip to content
AI-grafen
EUniversityData handling· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Personal data and anonymisation

Be able to identify personal data in a dataset and apply pseudonymisation and data minimisation.

Prerequisites

Intuition

Personal data is any information that can be linked to an identifiable natural person — directly or indirectly.

TypeExamples
Directly identifyingname, national ID number, email, phone number
Indirect (quasi-identifiers)postcode + date of birth + sex, IP address, device id
Special categorieshealth, ethnicity, religion, political opinion, sexual orientation, biometrics

The middle row is the one that surprises people. Sweeney showed that a postcode, a date of birth and a sex are enough to uniquely identify roughly 87 % of the US population. Removing the name is therefore nowhere near enough.

Formal

The distinction that decides what the law requires:

  • Pseudonymisation: the identifiers are replaced with an id, but a key still exists (or reidentification is possible). The data is still personal data and the GDPR applies in full.
  • Anonymisation: reidentification is impossible even with extra information. The GDPR then does not apply — but the threshold is high, and most «anonymised» datasets are in practice pseudonymised.

The techniques and what they actually give:

TechniqueProtectionLimitation
Removing the direct identifierslowthe quasi-identifiers remain
k-anonymityevery row is shared by at least k othersvulnerable to homogeneity and background-knowledge attacks
l-diversity, t-closenessvariation in the sensitive attributescomplex, destroys utility
Differential privacya mathematical guarantee via calibrated noisecosts accuracy; ε has to be chosen and justified
Aggregationhigh if the groups are largethe detail is lost

Data minimisation is the strongest and cheapest measure: do not collect what you do not need. Data that was never collected cannot leak, cannot be misused and does not have to be deleted.

For AI systems there is the additional fact that the model can memorise training data. Carlini et al. have shown that exact text strings can be extracted from language models. Personal data in the training data is therefore a risk even after the data has been deleted.

Code

import re, hashlib, os

PII = {
    "national_id": re.compile(r"\b(19|20)?\d{6}[-+]?\d{4}\b"),
    "email": re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+"),
    "phone": re.compile(r"\b0\d{1,3}[-\s]?\d{5,8}\b"),
    "card_number": re.compile(r"\b(?:\d[ -]?){13,16}\b"),
}

def mask(text: str) -> tuple[str, dict]:
    found = {}
    for name, m in PII.items():
        text, n = m.subn(f"<{name.upper()}>", text)
        if n:
            found[name] = n
    return text, found

PEPPER = os.environ["PSEUDONYM_PEPPER"].encode()      # secret, not rotated lightly
def pseudonymise(identifier: str) -> str:
    """A stable id without revealing the original. NOTE: still personal data."""
    return hashlib.blake2b(identifier.encode() + PEPPER, digest_size=8).hexdigest()

text = "Kontakta Anna på [email protected] eller 070-1234567."
print(mask(text))
# ('Kontakta Anna på <EMAIL> eller <PHONE>.', {'email': 1, 'phone': 1})
#            ↑ the name is not caught by the regex — that needs NER or manual review

The example shows the limitation: regexes catch structured identifiers but not names, places or free text that gives somebody away. Combine them with NER and human review of a sample.

Mastery means

  • Identifies direct and indirect personal data
  • Applies pseudonymisation and data minimisation
  • Understands why anonymisation is harder than it sounds

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences