Skip to content
AI-grafen
EUniversityAudio and speech· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Fourier and spectrograms

Be able to compute a spectrogram and explain why it is the input to most audio models.

Prerequisites

Intuition

A waveform is pressure over time: 16 000 numbers a second saying how the speaker cone should move. That says almost nothing about what the sound is.

A Fourier transform answers a different question: which frequencies does the sound consist of? It decomposes the signal into pure tones and says how much there is of each.

But a whole audio clip at once is too blunt — music and speech change all the time. So you do an STFT (short-time Fourier transform):

  1. Cut the signal into short, overlapping pieces (25 ms with a 10 ms hop, say).
  2. Fourier transform each piece separately.
  3. Lay the results side by side.

The result is a spectrogram: an image where the x axis is time, the y axis is frequency and the brightness is energy. That image is what nearly all audio models actually get to see.

Formal

The discrete Fourier transform of a frame x0..xN−1x_0..x_{N-1}:

Xk=∑n=0N−1xne−2πikn/N,k=0..N−1X_k = \sum_{n=0}^{N-1} x_n e^{-2\pi i kn/N}, \qquad k = 0..N-1

∣Xk∣|X_k| is the amplitude at the frequency fk=k⋅fs/Nf_k = k \cdot f_s / N hertz. For a real input the spectrum is symmetric, so only N/2+1N/2 + 1 values carry information — which is what rfft returns.

The time–frequency trade-off. With a sample rate fsf_s and a window length NN:

  • the frequency resolution is Δf=fs/N\Delta f = f_s / N,
  • the time resolution is N/fsN / f_s seconds.

They go in opposite directions. At 16 kHz:

WindowΔfTime resolutionSuits
128 (8 ms)125 Hz8 mstransients, drum hits
400 (25 ms)40 Hz25 msspeech (the standard choice)
2048 (128 ms)7.8 Hz128 mspitch in music

The window function. Simply cutting out a piece means multiplying by a rectangle, which spreads energy across the whole spectrum (spectral leakage). So every frame is multiplied by a smoothly tapering function — Hann is the standard choice.

Nyquist: frequencies above fs/2f_s/2 cannot be represented. Sampling at 16 kHz gives information up to 8 kHz, which is plenty for speech.

Code

import numpy as np

fs = 16000
t = np.arange(fs) / fs
# Two tones that swap after half a second
x = np.concatenate([np.sin(2 * np.pi * 440 * t[:fs // 2]),
                    np.sin(2 * np.pi * 880 * t[fs // 2:])])

def stft(x, N=400, hop=160):
    w = np.hanning(N)
    frames = 1 + (len(x) - N) // hop
    return np.stack([np.abs(np.fft.rfft(x[i * hop:i * hop + N] * w)) for i in range(frames)])

S = stft(x)
print(S.shape)                                  # (98, 201)
freq = np.fft.rfftfreq(400, 1 / fs)             # 0 … 8000 Hz in 201 steps
print(round(freq[np.argmax(S[10])]), round(freq[np.argmax(S[80])]))   # 440 880

# The dB scale — nearly always what you want to see and to feed in
S_db = 20 * np.log10(S + 1e-10)
print(round(S_db.max() - S_db.min()))           # the dynamic range in dB

The frame at index 10 lies in the first half (440 Hz), the frame at index 80 in the second (880 Hz) — the spectrogram shows the swap, which a single Fourier transform over the whole clip would not have.

The dB scale is not cosmetics. Hearing is roughly logarithmic, and a linear amplitude makes everything but the loudest parts go black. Models also train considerably better on a log scale.

Interactive

Try the time–frequency trade-off yourself on a clip of somebody saying a short sentence:

  1. Run the STFT with N=128, N=512 and N=2048 (hop = N/4).
  2. Plot all three as images with the same axes.

You should see:

  • N=128: every consonant shows up sharply in time, but the frequency bands run together into blurry blocks.
  • N=512: a balance — you see both the syllable boundaries and the horizontal stripes (the formants) that distinguish vowels.
  • N=2048: beautifully sharp frequency lines, but the syllables are smeared out and you can no longer see where a word ends.

The question to answer: which of the images would you feed into a model that has to decide which word is being said, and which into a model that has to decide which note a singer is holding? The answers differ, and that is the whole point.

Mastery means

  • Explains what a Fourier transform does to a signal
  • Computes a spectrogram with the STFT
  • Chooses the window length from the time–frequency trade-off

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences