Speech recognition (ASR)
Be able to explain CTC and encoder–decoder ASR and run Whisper-like models.
Prerequisites
- DTransformers — the architecturerequired
- EThe mel scale and MFCCsrequired
Intuition
The basic problem in ASR: the audio has thousands of frames, the transcription has tens of characters, and nobody says where the boundaries are. Two families of solutions:
CTC. Let the model guess one character per frame, with an extra «blank» character. Then merge repetitions and remove the blanks: h h _ i i → hi. The training sums over all the paths that give the right output. Fast, streaming, but every frame is decided independently — no language modelling built in.
Encoder–decoder. The encoder encodes the whole audio, the decoder generates the text token by token and gets to look at the audio through cross-attention. Better linguistic quality, but not streaming, and it can hallucinate text that was never said.
Whisper is an encoder–decoder, trained on 680 000 hours of weakly labelled audio from the web — and it is that scale, not the architecture, that explains why it is robust to noise and accents.
Formal
The CTC loss. For audio frames and a target text with : let be the function that merges repetitions and removes blanks. Then
The sum over exponentially many paths is computed in with the forward–backward algorithm — the same dynamic programming as in hidden Markov models.
WER (word error rate) is the edit distance at the word level:
where , and are substitutions, deletions and insertions and is the number of words in the reference. WER can exceed 100 % (many insertions).
Normalise before you measure — otherwise you are measuring formatting, not recognition: lower case, remove punctuation, numbers to words or the other way round, consistently. An implausibly high WER is more often because «5» is being compared with «five» than because of the model.
Swedish in practice:
| Choice | Effect |
|---|---|
language="sv" explicitly | stops the model guessing the language and transcribing into English |
| KB-Whisper instead of a generic model | trained on Swedish material, a noticeably lower WER |
| VAD segmentation first | strongly reduces hallucinations in silent passages |
condition_on_previous_text=False | breaks loops where the model repeats the same phrase |
The third row is the most important in practice: Whisper models hallucinate above all in silence, and a simple voice activity detector before the transcription removes most of it.
Code
import re, jiwer
from transformers import pipeline
asr = pipeline("automatic-speech-recognition", model="KBLab/kb-whisper-small",
chunk_length_s=30, return_timestamps=True)
out = asr("meeting.wav", generate_kwargs={"language": "sv", "task": "transcribe",
"condition_on_previous_text": False})
NORM = jiwer.Compose([jiwer.ToLowerCase(), jiwer.RemovePunctuation(),
jiwer.RemoveMultipleSpaces(), jiwer.Strip(),
jiwer.ReduceToListOfListOfWords()])
def wer(reference, hypothesis):
return jiwer.wer(reference, hypothesis, truth_transform=NORM, hypothesis_transform=NORM)
print(round(wer("Vi ses klockan tre på torsdag", out["text"]), 3))
# A hallucination detector: the same phrase repeated in sequence
def suspicious(text, n=3):
words = text.lower().split()
for i in range(len(words) - 2 * n):
if words[i:i + n] == words[i + n:i + 2 * n]:
return True
return False
suspicious is invaluable in production: it catches the classic Whisper fault where a silent passage produces «thanks for watching» over and over. Combine it with the timestamps — a segment with an unusually high number of characters per second is nearly always a hallucination.
Mastery means
- Explains the difference between CTC and encoder–decoder
- Measures WER correctly
- Chooses the model and the settings for Swedish
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Robust Speech Recognition via Large-Scale Weak Supervision — arXiv (open access; licence per article)
- Kungliga biblioteket — KB-Whisper — open models
- Graves m.fl. — Connectionist Temporal Classification (ICML 2006) — author's copy, free to read