Skip to content
AI-grafen
EUniversityAudio and speech· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

The mel scale and MFCCs

Be able to compute mel spectrograms and explain why the scale imitates hearing.

Prerequisites

Intuition

Hearing is not linear in frequency. The difference between 100 and 200 Hz sounds like a big step; the difference between 8 000 and 8 100 Hz you can barely hear at all. Both are 100 Hz.

The mel scale is a rescaling that makes equal steps sound equal. It is nearly linear below 1 000 Hz and logarithmic above.

A mel spectrogram is built in three steps:

  1. STFT → a magnitude spectrogram (201 frequency bins at a 25 ms window and 16 kHz).
  2. Sum the bins together with triangular filters that are narrow at the bottom and wide at the top — typically 80 filters.
  3. Take the logarithm.

The result is more compact (80 channels instead of 201), more perceptually relevant, and it is what Whisper and nearly all modern audio models take as input.

Formal

The mel conversion (the most common variant):

m=2595log⁡10 ⁣(1+f700),f=700(10m/2595−1)m = 2595 \log_{10}\!\left(1 + \frac{f}{700}\right), \qquad f = 700\left(10^{m/2595} - 1\right)

The scale is calibrated so that 1 000 Hz ≈ 1 000 mel.

The filterbank. Place M+2M + 2 points evenly spaced on the mel scale between fmin⁡f_{\min} and fmax⁡f_{\max}, convert them back to hertz, and build triangular filters where filter mm has its support between point mm and m+2m+2 with its peak at m+1m+1. Even spacing in mel is therefore uneven in hertz: the filters are narrow at the bottom and wide at the top. That is the whole point.

MFCCs go one step further: take the DCT of the log-mel energies and keep the first 13 coefficients. The DCT decorrelates the channels, which was necessary for the Gaussian mixture models that dominated until about 2012.

When do you need what?

MethodChannelsUsed for
Log-mel spectrogram40–128neural networks — the standard today
MFCC13 (+Δ, ΔΔ)classical ML, small datasets, very constrained hardware
Raw waveform—wav2vec 2.0 and similar self-supervised models

A practical rule: use log-mel if you are training a neural network. MFCCs are not wrong, but the DCT throws away information a CNN would rather have used itself. Convolutional networks handle correlated channels perfectly well — they are built for it.

Code

import numpy as np

def hz2mel(f): return 2595.0 * np.log10(1.0 + np.asarray(f, float) / 700.0)
def mel2hz(m): return 700.0 * (10.0 ** (np.asarray(m, float) / 2595.0) - 1.0)

def mel_filterbank(n_filters=80, n_fft=400, fs=16000, f_min=0.0, f_max=8000.0):
    points = mel2hz(np.linspace(hz2mel(f_min), hz2mel(f_max), n_filters + 2))
    bins = np.floor((n_fft + 1) * points / fs).astype(int)
    fb = np.zeros((n_filters, n_fft // 2 + 1))
    for m in range(n_filters):
        left, top, right = bins[m], bins[m + 1], bins[m + 2]
        for k in range(left, top):
            fb[m, k] = (k - left) / max(top - left, 1)
        for k in range(top, right):
            fb[m, k] = (right - k) / max(right - top, 1)
    return fb

fb = mel_filterbank()
width = fb.sum(axis=1)
print(round(width[0], 1), round(width[-1], 1))   # ~1.0  ~10.5 — the filters get wider upwards

# a log-mel spectrogram from a magnitude spectrogram S (frames × bins)
mel = np.log(fb @ S.T + 1e-10).T              # (frames, 80)

The line width[0] against width[-1] shows the central point: the topmost filter covers ten times more frequency bins than the bottom one. The model therefore gets high resolution where the ear has it, and coarse resolution where the ear does not care.

Mastery means

  • Converts between hertz and mel
  • Builds a mel filterbank
  • Knows when MFCCs are needed and when a mel spectrogram is enough

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences