Fourier and spectrograms
Be able to compute a spectrogram and explain why it is the input to most audio models.
Prerequisites
Intuition
A waveform is pressure over time: 16 000 numbers a second saying how the speaker cone should move. That says almost nothing about what the sound is.
A Fourier transform answers a different question: which frequencies does the sound consist of? It decomposes the signal into pure tones and says how much there is of each.
But a whole audio clip at once is too blunt — music and speech change all the time. So you do an STFT (short-time Fourier transform):
- Cut the signal into short, overlapping pieces (25 ms with a 10 ms hop, say).
- Fourier transform each piece separately.
- Lay the results side by side.
The result is a spectrogram: an image where the x axis is time, the y axis is frequency and the brightness is energy. That image is what nearly all audio models actually get to see.
Formal
The discrete Fourier transform of a frame :
is the amplitude at the frequency hertz. For a real input the spectrum is symmetric, so only values carry information — which is what rfft returns.
The time–frequency trade-off. With a sample rate and a window length :
- the frequency resolution is ,
- the time resolution is seconds.
They go in opposite directions. At 16 kHz:
| Window | Δf | Time resolution | Suits |
|---|---|---|---|
| 128 (8 ms) | 125 Hz | 8 ms | transients, drum hits |
| 400 (25 ms) | 40 Hz | 25 ms | speech (the standard choice) |
| 2048 (128 ms) | 7.8 Hz | 128 ms | pitch in music |
The window function. Simply cutting out a piece means multiplying by a rectangle, which spreads energy across the whole spectrum (spectral leakage). So every frame is multiplied by a smoothly tapering function — Hann is the standard choice.
Nyquist: frequencies above cannot be represented. Sampling at 16 kHz gives information up to 8 kHz, which is plenty for speech.
Code
import numpy as np
fs = 16000
t = np.arange(fs) / fs
# Two tones that swap after half a second
x = np.concatenate([np.sin(2 * np.pi * 440 * t[:fs // 2]),
np.sin(2 * np.pi * 880 * t[fs // 2:])])
def stft(x, N=400, hop=160):
w = np.hanning(N)
frames = 1 + (len(x) - N) // hop
return np.stack([np.abs(np.fft.rfft(x[i * hop:i * hop + N] * w)) for i in range(frames)])
S = stft(x)
print(S.shape) # (98, 201)
freq = np.fft.rfftfreq(400, 1 / fs) # 0 … 8000 Hz in 201 steps
print(round(freq[np.argmax(S[10])]), round(freq[np.argmax(S[80])])) # 440 880
# The dB scale — nearly always what you want to see and to feed in
S_db = 20 * np.log10(S + 1e-10)
print(round(S_db.max() - S_db.min())) # the dynamic range in dB
The frame at index 10 lies in the first half (440 Hz), the frame at index 80 in the second (880 Hz) — the spectrogram shows the swap, which a single Fourier transform over the whole clip would not have.
The dB scale is not cosmetics. Hearing is roughly logarithmic, and a linear amplitude makes everything but the loudest parts go black. Models also train considerably better on a log scale.
Interactive
Try the time–frequency trade-off yourself on a clip of somebody saying a short sentence:
- Run the STFT with
N=128,N=512andN=2048(hop = N/4). - Plot all three as images with the same axes.
You should see:
- N=128: every consonant shows up sharply in time, but the frequency bands run together into blurry blocks.
- N=512: a balance — you see both the syllable boundaries and the horizontal stripes (the formants) that distinguish vowels.
- N=2048: beautifully sharp frequency lines, but the syllables are smeared out and you can no longer see where a word ends.
The question to answer: which of the images would you feed into a model that has to decide which word is being said, and which into a model that has to decide which note a singer is holding? The answers differ, and that is the whole point.
Mastery means
- Explains what a Fourier transform does to a signal
- Computes a spectrogram with the STFT
- Chooses the window length from the time–frequency trade-off
Sign in to do the exercises and build your mastery up.
Sources
- librosa — dokumentation (ISC) — ISC
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0
- Wikipedia — Short-time Fourier transform (CC BY-SA 4.0) — CC BY-SA 4.0