Skip to content
AI-grafen
EUniversityAudio and speech· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Audio classification

Be able to train a classifier on spectrograms.

Prerequisites

Intuition

A mel spectrogram is an image. Which is why an ordinary CNN works surprisingly well on audio — the whole toolbox from image classification can be reused.

But three things differ from images:

  1. The axes mean different things. Moving a pattern in time changes nothing (the same sound, later). Moving it in frequency changes everything (a different pitch). Translation invariance is desirable in one direction but not in the other.
  2. The augmentation is audio-specific. Time stretching, pitch shifting, background noise, and SpecAugment — masking random bands in time and frequency directly in the spectrogram.
  3. Leakage is easier to fall into. Cut a long clip into pieces and scatter them randomly across the training and test sets, and you have the same recording on both sides. The model learns the room, the microphone or the speaker — not the class.

Code

import torch, torch.nn as nn, torchaudio

mel = torchaudio.transforms.MelSpectrogram(sample_rate=16000, n_fft=400,
                                           hop_length=160, n_mels=64)
to_db = torchaudio.transforms.AmplitudeToDB()

# SpecAugment: mask bands in time and frequency
aug = nn.Sequential(torchaudio.transforms.FrequencyMasking(freq_mask_param=12),
                    torchaudio.transforms.TimeMasking(time_mask_param=25))

class AudioCNN(nn.Module):
    def __init__(self, n_classes):
        super().__init__()
        def block(i, o):
            return nn.Sequential(nn.Conv2d(i, o, 3, padding=1), nn.BatchNorm2d(o),
                                 nn.ReLU(), nn.MaxPool2d(2))
        self.f = nn.Sequential(block(1, 32), block(32, 64), block(64, 128),
                               nn.AdaptiveAvgPool2d(1), nn.Flatten())
        self.h = nn.Linear(128, n_classes)

    def forward(self, waveform):
        x = to_db(mel(waveform)).unsqueeze(1)           # (B, 1, mel, time)
        x = (x - x.mean()) / (x.std() + 1e-5)
        if self.training:
            x = aug(x)
        return self.h(self.f(x))

Split by recording, not by clip. Make the split at the recording level (or the speaker level) before you cut:

import numpy as np

recordings = sorted({c["recording_id"] for c in clips})
rng = np.random.default_rng(0); rng.shuffle(recordings)
n = int(0.8 * len(recordings))
train_ids, test_ids = set(recordings[:n]), set(recordings[n:])
train = [c for c in clips if c["recording_id"] in train_ids]
test  = [c for c in clips if c["recording_id"] in test_ids]

The difference between the right and the wrong split is often 15–25 percentage points of reported accuracy — pure illusion.

Interactive

A real failure pattern. A model classifying bird calls reaches 94 % on validation and 61 % in the field. The debugging usually goes like this:

CheckWhat you are looking for
Was the data split by recording?the same background noise in training and test
Does the class correlate with the recording location?the model learns the place, not the bird
Do the microphone or the sample rate differ?domain shift between the lab and the field
What does a model that sees only the background noise give?if it reaches 70 % the shortcut is there

The last check is the most revealing and the one most often skipped: train a model on clips where you have masked out the call itself and kept the background. If it beats chance you have a shortcut in the data, not a classifier.

The same logic applies to all audio: a «cough detector» that really recognises hospital rooms, an «engine fault detector» that recognises the workshop.

Mastery means

  • Trains a CNN on mel spectrograms
  • Applies audio-specific augmentation
  • Avoids leakage between recordings

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences