Voice data and privacy
Be able to reason about biometric voice data, consent and the risks of voice cloning.
Prerequisites
Intuition
A voice is not just a way of carrying words. From a recording you can read off who is speaking, their approximate age and sex, their dialect and first language, their emotional state, and in some cases indications of health.
That makes voice data something other than text. And unlike a password, you cannot change your voice once it has leaked.
Two distinct things that are often confused:
- Speaker recognition — who is speaking? (biometrics)
- Speech recognition — what is being said? (transcription)
The first is always sensitive. The second becomes sensitive through its content.
Formal
The legal position in the EU. Voice data used to identify a person is biometric data and belongs to the special categories in GDPR Article 9. Processing is as a rule prohibited unless an exception applies — in practice most often explicit consent, which has to be freely given, specific, informed and as easy to withdraw as to give.
If the recording is used only to transcribe what is said it is «ordinary» personal data (Article 6) — but the content can still be sensitive.
The EU AI Act adds: real-time remote identification in public spaces via biometrics is heavily restricted, and emotion recognition in workplaces and in education is prohibited. That last point is directly relevant to learning platforms — a feature that judges a pupil's emotional state from their voice is not a design question but not permitted.
Safeguards that actually make a difference:
| Measure | Effect |
|---|---|
| Transcribe and delete the audio | removes the biometrics entirely |
| Store an embedding instead of the audio | less reconstructable, but still biometric data |
| Edge transcription (on the device) | the audio never leaves the device |
| A short retention time with automatic purging | reduces the exposure |
| Separate keys and strict access control | limits the damage in a breach |
Voice cloning. A few seconds of audio is enough today to generate convincing speech in somebody's voice. The consequences are concrete: fraud against relatives, circumventing voice-based authentication (which should therefore no longer be used alone), defamation and harassment. A voice as the sole authentication factor is no longer defensible.
The countermeasures resemble those for images: watermarking and provenance can establish authenticity, but the absence of a mark proves nothing. The effective protection is procedural — calling back on a known number, an agreed code word within the family, two-factor authentication that is not built on a voice.
Interactive
Assess a real design. AI-grafen could let pupils practise oral presentation by recording themselves and getting feedback.
Work through the questions in order:
- What is the purpose? Feedback on content and structure — not identifying who is speaking.
- Is the audio needed after the transcription? No, for content feedback; yes, if pronunciation is to be assessed.
- The legal basis? Explicit consent; for children under 13, the guardian's.
- Storage? The transcript is kept, the audio is deleted after processing — or kept for a maximum of 30 days if pronunciation is to be assessed, with clear information.
- What is prohibited? Emotion recognition in an educational setting — even if it would be technically easy and «only» used to detect nervousness.
- What happens if they decline? The pupil must be able to do the same task without recording.
Point 6 is what decides whether the consent is freely given. If there is no way round it, it is not consent.
Mastery means
- Explains why the voice is biometric personal data
- Applies the right legal basis and safeguards
- Assesses the risks of voice cloning realistically
Sign in to do the exercises and build your mastery up.
Sources
- GDPR — Regulation (EU) 2016/679, Art. 9 — EU legal act
- EU AI Act (2024/1689), Art. 5 — EU legal act
- IMY — biometrics and privacy (in Swedish) — myndighetsmaterial