Speech synthesis (TTS)
Be able to explain text-to-mel and the vocoder and to evaluate synthesised speech.
Prerequisites
- EGenerative models — an overviewrequired
- EThe mel scale and MFCCsrequired
Intuition
Classic TTS has two steps:
- The acoustic model: text → a mel spectrogram. This is where pronunciation, stress, rhythm and melody are decided.
- The vocoder: mel spectrogram → a waveform. This is where it is decided whether it sounds natural or robotic.
The split exists because the two problems are so different. Step 1 is about language, step 2 about signal processing. Newer systems do everything in one step (VITS, and LLM-like models that generate audio tokens directly), but the two-step picture is still the best mental model.
The hard part is not sounding human — it is the prosody. The same sentence can mean entirely different things depending on where the stress falls. A TTS that pronounces every word correctly but stresses the wrong one sounds odd in a way that is hard to put your finger on.
In Swedish there is a particular extra difficulty: accent 1 and accent 2. «Anden» (the duck) and «anden» (the spirit) are distinguished only by the pitch accent. There is nothing in the spelling that decides it.
Formal
The chain in detail:
| Step | What happens | Typical models |
|---|---|---|
| Text normalisation | «kl. 14:30» → «klockan fjorton och trettio» | rules + word lists |
| Grapheme→phoneme | spelling → pronunciation, with an exception lexicon | a lexicon + a model |
| The acoustic model | phonemes + durations → mel | FastSpeech 2, Tacotron 2 |
| The vocoder | mel → a waveform | HiFi-GAN, WaveGlow |
Text normalisation is where most of the audible errors arise, not in the neural network. Abbreviations, numbers, dates, units, proper nouns and foreign words. «3 m/s» should become «three metres per second», not «three m slash s».
Evaluation:
| Method | Measures | Comment |
|---|---|---|
| MOS | perceived quality, 1–5, by listeners | the gold standard, expensive |
| MUSHRA / AB tests | a comparison between systems | more sensitive than MOS |
| UTMOS and similar | a predicted MOS | cheap, for regression tests |
| WER via ASR | intelligibility | catches pronunciation errors automatically |
| A pronunciation list | specific words | proper nouns, terms, abbreviations |
The fourth row is underrated: run your TTS through an ASR model and measure the WER against the text you fed in. That detects pronunciation errors automatically, at scale, and without listeners.
Voice cloning and consent. Cloning a voice today takes very little audio. The minimum requirements for responsible use: documented consent from the voice, clear labelling of synthetic speech (a requirement in the EU AI Act), traceability over who generated what, and a blocklist against known voices. Commercial providers that do this properly require a verification recording where the speaker reads a given text — not just an uploaded file.
Code
import re, torch, jiwer
from transformers import pipeline
# 1. Text normalisation — this is where most of the audible errors arise
UNITS = {"m/s": "meter per sekund", "kr": "kronor", "kl.": "klockan", "%": "procent"}
def normalise(t: str) -> str:
for k, v in UNITS.items():
t = t.replace(k, " " + v + " ")
t = re.sub(r"\b(\d{1,2}):(\d{2})\b", r"\1 och \2", t)
return re.sub(r"\s+", " ", t).strip()
print(normalise("Mötet börjar kl. 14:30 och det blåser 3 m/s."))
# Mötet börjar klockan 14 och 30 och det blåser 3 meter per sekund .
# ^ the space before the full stop is exactly the kind of error you can hear — clean up after normalising
# 2. Synthesise and 3. measure intelligibility automatically with ASR
tts = pipeline("text-to-speech", model="facebook/mms-tts-swe")
asr = pipeline("automatic-speech-recognition", model="KBLab/kb-whisper-small")
def intelligibility(sentences):
errors = []
for s in sentences:
audio = tts(normalise(s))
back = asr({"raw": audio["audio"].squeeze(), "sampling_rate": audio["sampling_rate"]},
generate_kwargs={"language": "sv"})["text"]
w = jiwer.wer(s.lower(), back.lower())
if w > 0.1:
errors.append((s, back, round(w, 2)))
return errors
for row in intelligibility(["Anden simmade i dammen.", "Tåget avgår kl. 14:30 från spår 3."]):
print(row)
That loop is a perfectly good regression suite for TTS: run it on every model or lexicon change and review only the sentences where the WER rises.
Mastery means
- Describes the TTS chain from text to waveform
- Evaluates synthesised speech with the right methods
- Handles voice cloning responsibly
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — FastSpeech 2: Fast and High-Quality End-to-End Text to Speech — arXiv (open access; licence per article)
- arXiv — HiFi-GAN — arXiv (open access; licence per article)
- EU AI Act (2024/1689), Art. 50 — EU legal act