Skip to content
AI-grafen
FAI engineeringAudio and speech· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Speech synthesis (TTS)

Be able to explain text-to-mel and the vocoder and to evaluate synthesised speech.

Prerequisites

Intuition

Classic TTS has two steps:

  1. The acoustic model: text → a mel spectrogram. This is where pronunciation, stress, rhythm and melody are decided.
  2. The vocoder: mel spectrogram → a waveform. This is where it is decided whether it sounds natural or robotic.

The split exists because the two problems are so different. Step 1 is about language, step 2 about signal processing. Newer systems do everything in one step (VITS, and LLM-like models that generate audio tokens directly), but the two-step picture is still the best mental model.

The hard part is not sounding human — it is the prosody. The same sentence can mean entirely different things depending on where the stress falls. A TTS that pronounces every word correctly but stresses the wrong one sounds odd in a way that is hard to put your finger on.

In Swedish there is a particular extra difficulty: accent 1 and accent 2. «Anden» (the duck) and «anden» (the spirit) are distinguished only by the pitch accent. There is nothing in the spelling that decides it.

Formal

The chain in detail:

StepWhat happensTypical models
Text normalisation«kl. 14:30» → «klockan fjorton och trettio»rules + word lists
Grapheme→phonemespelling → pronunciation, with an exception lexicona lexicon + a model
The acoustic modelphonemes + durations → melFastSpeech 2, Tacotron 2
The vocodermel → a waveformHiFi-GAN, WaveGlow

Text normalisation is where most of the audible errors arise, not in the neural network. Abbreviations, numbers, dates, units, proper nouns and foreign words. «3 m/s» should become «three metres per second», not «three m slash s».

Evaluation:

MethodMeasuresComment
MOSperceived quality, 1–5, by listenersthe gold standard, expensive
MUSHRA / AB testsa comparison between systemsmore sensitive than MOS
UTMOS and similara predicted MOScheap, for regression tests
WER via ASRintelligibilitycatches pronunciation errors automatically
A pronunciation listspecific wordsproper nouns, terms, abbreviations

The fourth row is underrated: run your TTS through an ASR model and measure the WER against the text you fed in. That detects pronunciation errors automatically, at scale, and without listeners.

Voice cloning and consent. Cloning a voice today takes very little audio. The minimum requirements for responsible use: documented consent from the voice, clear labelling of synthetic speech (a requirement in the EU AI Act), traceability over who generated what, and a blocklist against known voices. Commercial providers that do this properly require a verification recording where the speaker reads a given text — not just an uploaded file.

Code

import re, torch, jiwer
from transformers import pipeline

# 1. Text normalisation — this is where most of the audible errors arise
UNITS = {"m/s": "meter per sekund", "kr": "kronor", "kl.": "klockan", "%": "procent"}

def normalise(t: str) -> str:
    for k, v in UNITS.items():
        t = t.replace(k, " " + v + " ")
    t = re.sub(r"\b(\d{1,2}):(\d{2})\b", r"\1 och \2", t)
    return re.sub(r"\s+", " ", t).strip()

print(normalise("Mötet börjar kl. 14:30 och det blåser 3 m/s."))
# Mötet börjar klockan 14 och 30 och det blåser 3 meter per sekund .
# ^ the space before the full stop is exactly the kind of error you can hear — clean up after normalising

# 2. Synthesise and 3. measure intelligibility automatically with ASR
tts = pipeline("text-to-speech", model="facebook/mms-tts-swe")
asr = pipeline("automatic-speech-recognition", model="KBLab/kb-whisper-small")

def intelligibility(sentences):
    errors = []
    for s in sentences:
        audio = tts(normalise(s))
        back = asr({"raw": audio["audio"].squeeze(), "sampling_rate": audio["sampling_rate"]},
                   generate_kwargs={"language": "sv"})["text"]
        w = jiwer.wer(s.lower(), back.lower())
        if w > 0.1:
            errors.append((s, back, round(w, 2)))
    return errors

for row in intelligibility(["Anden simmade i dammen.", "Tåget avgår kl. 14:30 från spår 3."]):
    print(row)

That loop is a perfectly good regression suite for TTS: run it on every model or lexicon change and review only the sentences where the WER rises.

Mastery means

  • Describes the TTS chain from text to waveform
  • Evaluates synthesised speech with the right methods
  • Handles voice cloning responsibly

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences