Audio classification
Be able to train a classifier on spectrograms.
Prerequisites
- EConvolutional networks (CNNs)required
- EFourier and spectrogramsrequired
Intuition
A mel spectrogram is an image. Which is why an ordinary CNN works surprisingly well on audio — the whole toolbox from image classification can be reused.
But three things differ from images:
- The axes mean different things. Moving a pattern in time changes nothing (the same sound, later). Moving it in frequency changes everything (a different pitch). Translation invariance is desirable in one direction but not in the other.
- The augmentation is audio-specific. Time stretching, pitch shifting, background noise, and SpecAugment — masking random bands in time and frequency directly in the spectrogram.
- Leakage is easier to fall into. Cut a long clip into pieces and scatter them randomly across the training and test sets, and you have the same recording on both sides. The model learns the room, the microphone or the speaker — not the class.
Code
import torch, torch.nn as nn, torchaudio
mel = torchaudio.transforms.MelSpectrogram(sample_rate=16000, n_fft=400,
hop_length=160, n_mels=64)
to_db = torchaudio.transforms.AmplitudeToDB()
# SpecAugment: mask bands in time and frequency
aug = nn.Sequential(torchaudio.transforms.FrequencyMasking(freq_mask_param=12),
torchaudio.transforms.TimeMasking(time_mask_param=25))
class AudioCNN(nn.Module):
def __init__(self, n_classes):
super().__init__()
def block(i, o):
return nn.Sequential(nn.Conv2d(i, o, 3, padding=1), nn.BatchNorm2d(o),
nn.ReLU(), nn.MaxPool2d(2))
self.f = nn.Sequential(block(1, 32), block(32, 64), block(64, 128),
nn.AdaptiveAvgPool2d(1), nn.Flatten())
self.h = nn.Linear(128, n_classes)
def forward(self, waveform):
x = to_db(mel(waveform)).unsqueeze(1) # (B, 1, mel, time)
x = (x - x.mean()) / (x.std() + 1e-5)
if self.training:
x = aug(x)
return self.h(self.f(x))
Split by recording, not by clip. Make the split at the recording level (or the speaker level) before you cut:
import numpy as np
recordings = sorted({c["recording_id"] for c in clips})
rng = np.random.default_rng(0); rng.shuffle(recordings)
n = int(0.8 * len(recordings))
train_ids, test_ids = set(recordings[:n]), set(recordings[n:])
train = [c for c in clips if c["recording_id"] in train_ids]
test = [c for c in clips if c["recording_id"] in test_ids]
The difference between the right and the wrong split is often 15–25 percentage points of reported accuracy — pure illusion.
Interactive
A real failure pattern. A model classifying bird calls reaches 94 % on validation and 61 % in the field. The debugging usually goes like this:
| Check | What you are looking for |
|---|---|
| Was the data split by recording? | the same background noise in training and test |
| Does the class correlate with the recording location? | the model learns the place, not the bird |
| Do the microphone or the sample rate differ? | domain shift between the lab and the field |
| What does a model that sees only the background noise give? | if it reaches 70 % the shortcut is there |
The last check is the most revealing and the one most often skipped: train a model on clips where you have masked out the call itself and kept the background. If it beats chance you have a shortcut in the data, not a classifier.
The same logic applies to all audio: a «cough detector» that really recognises hospital rooms, an «engine fault detector» that recognises the workshop.
Mastery means
- Trains a CNN on mel spectrograms
- Applies audio-specific augmentation
- Avoids leakage between recordings
Sign in to do the exercises and build your mastery up.
Sources
- torchaudio — dokumentation (BSD-2) — BSD-2-Clause
- arXiv — SpecAugment — arXiv (open access; licence per article)
- librosa — dokumentation (ISC) — ISC