Skip to content
AI-grafen
FAI engineeringAudio and speech· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Speech recognition (ASR)

Be able to explain CTC and encoder–decoder ASR and run Whisper-like models.

Prerequisites

Intuition

The basic problem in ASR: the audio has thousands of frames, the transcription has tens of characters, and nobody says where the boundaries are. Two families of solutions:

CTC. Let the model guess one character per frame, with an extra «blank» character. Then merge repetitions and remove the blanks: h h _ i i → hi. The training sums over all the paths that give the right output. Fast, streaming, but every frame is decided independently — no language modelling built in.

Encoder–decoder. The encoder encodes the whole audio, the decoder generates the text token by token and gets to look at the audio through cross-attention. Better linguistic quality, but not streaming, and it can hallucinate text that was never said.

Whisper is an encoder–decoder, trained on 680 000 hours of weakly labelled audio from the web — and it is that scale, not the architecture, that explains why it is robust to noise and accents.

Formal

The CTC loss. For audio frames x1:Tx_{1:T} and a target text y1:Uy_{1:U} with T≫UT \gg U: let B\mathcal{B} be the function that merges repetitions and removes blanks. Then

P(y∣x)=∑π∈B−1(y)∏t=1TP(πt∣x)P(y \mid x) = \sum_{\pi \in \mathcal{B}^{-1}(y)} \prod_{t=1}^{T} P(\pi_t \mid x)

The sum over exponentially many paths is computed in O(TU)O(TU) with the forward–backward algorithm — the same dynamic programming as in hidden Markov models.

WER (word error rate) is the edit distance at the word level:

WER=S+D+IN\mathrm{WER} = \frac{S + D + I}{N}

where SS, DD and II are substitutions, deletions and insertions and NN is the number of words in the reference. WER can exceed 100 % (many insertions).

Normalise before you measure — otherwise you are measuring formatting, not recognition: lower case, remove punctuation, numbers to words or the other way round, consistently. An implausibly high WER is more often because «5» is being compared with «five» than because of the model.

Swedish in practice:

ChoiceEffect
language="sv" explicitlystops the model guessing the language and transcribing into English
KB-Whisper instead of a generic modeltrained on Swedish material, a noticeably lower WER
VAD segmentation firststrongly reduces hallucinations in silent passages
condition_on_previous_text=Falsebreaks loops where the model repeats the same phrase

The third row is the most important in practice: Whisper models hallucinate above all in silence, and a simple voice activity detector before the transcription removes most of it.

Code

import re, jiwer
from transformers import pipeline

asr = pipeline("automatic-speech-recognition", model="KBLab/kb-whisper-small",
               chunk_length_s=30, return_timestamps=True)
out = asr("meeting.wav", generate_kwargs={"language": "sv", "task": "transcribe",
                                          "condition_on_previous_text": False})

NORM = jiwer.Compose([jiwer.ToLowerCase(), jiwer.RemovePunctuation(),
                      jiwer.RemoveMultipleSpaces(), jiwer.Strip(),
                      jiwer.ReduceToListOfListOfWords()])

def wer(reference, hypothesis):
    return jiwer.wer(reference, hypothesis, truth_transform=NORM, hypothesis_transform=NORM)

print(round(wer("Vi ses klockan tre på torsdag", out["text"]), 3))

# A hallucination detector: the same phrase repeated in sequence
def suspicious(text, n=3):
    words = text.lower().split()
    for i in range(len(words) - 2 * n):
        if words[i:i + n] == words[i + n:i + 2 * n]:
            return True
    return False

suspicious is invaluable in production: it catches the classic Whisper fault where a silent passage produces «thanks for watching» over and over. Combine it with the timestamps — a segment with an unusually high number of characters per second is nearly always a hallucination.

Mastery means

  • Explains the difference between CTC and encoder–decoder
  • Measures WER correctly
  • Chooses the model and the settings for Swedish

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences