The mel scale and MFCCs
Be able to compute mel spectrograms and explain why the scale imitates hearing.
Prerequisites
- EFourier and spectrogramsrequired
Intuition
Hearing is not linear in frequency. The difference between 100 and 200 Hz sounds like a big step; the difference between 8 000 and 8 100 Hz you can barely hear at all. Both are 100 Hz.
The mel scale is a rescaling that makes equal steps sound equal. It is nearly linear below 1 000 Hz and logarithmic above.
A mel spectrogram is built in three steps:
- STFT → a magnitude spectrogram (201 frequency bins at a 25 ms window and 16 kHz).
- Sum the bins together with triangular filters that are narrow at the bottom and wide at the top — typically 80 filters.
- Take the logarithm.
The result is more compact (80 channels instead of 201), more perceptually relevant, and it is what Whisper and nearly all modern audio models take as input.
Formal
The mel conversion (the most common variant):
The scale is calibrated so that 1 000 Hz ≈ 1 000 mel.
The filterbank. Place points evenly spaced on the mel scale between and , convert them back to hertz, and build triangular filters where filter has its support between point and with its peak at . Even spacing in mel is therefore uneven in hertz: the filters are narrow at the bottom and wide at the top. That is the whole point.
MFCCs go one step further: take the DCT of the log-mel energies and keep the first 13 coefficients. The DCT decorrelates the channels, which was necessary for the Gaussian mixture models that dominated until about 2012.
When do you need what?
| Method | Channels | Used for |
|---|---|---|
| Log-mel spectrogram | 40–128 | neural networks — the standard today |
| MFCC | 13 (+Δ, ΔΔ) | classical ML, small datasets, very constrained hardware |
| Raw waveform | — | wav2vec 2.0 and similar self-supervised models |
A practical rule: use log-mel if you are training a neural network. MFCCs are not wrong, but the DCT throws away information a CNN would rather have used itself. Convolutional networks handle correlated channels perfectly well — they are built for it.
Code
import numpy as np
def hz2mel(f): return 2595.0 * np.log10(1.0 + np.asarray(f, float) / 700.0)
def mel2hz(m): return 700.0 * (10.0 ** (np.asarray(m, float) / 2595.0) - 1.0)
def mel_filterbank(n_filters=80, n_fft=400, fs=16000, f_min=0.0, f_max=8000.0):
points = mel2hz(np.linspace(hz2mel(f_min), hz2mel(f_max), n_filters + 2))
bins = np.floor((n_fft + 1) * points / fs).astype(int)
fb = np.zeros((n_filters, n_fft // 2 + 1))
for m in range(n_filters):
left, top, right = bins[m], bins[m + 1], bins[m + 2]
for k in range(left, top):
fb[m, k] = (k - left) / max(top - left, 1)
for k in range(top, right):
fb[m, k] = (right - k) / max(right - top, 1)
return fb
fb = mel_filterbank()
width = fb.sum(axis=1)
print(round(width[0], 1), round(width[-1], 1)) # ~1.0 ~10.5 — the filters get wider upwards
# a log-mel spectrogram from a magnitude spectrogram S (frames × bins)
mel = np.log(fb @ S.T + 1e-10).T # (frames, 80)
The line width[0] against width[-1] shows the central point: the topmost filter covers ten times more frequency bins than the bottom one. The model therefore gets high resolution where the ear has it, and coarse resolution where the ear does not care.
Mastery means
- Converts between hertz and mel
- Builds a mel filterbank
- Knows when MFCCs are needed and when a mel spectrogram is enough
Sign in to do the exercises and build your mastery up.
Sources
- librosa — dokumentation (ISC) — ISC
- Wikipedia — Mel scale (CC BY-SA 4.0) — CC BY-SA 4.0
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0