Training a model for audio classification
Be able to train a model that distinguishes between two sounds and test it.
Prerequisites
Intuition
Just like with images, a model can learn to distinguish sounds: a click versus a snap, a 'yes' versus a 'no', or one specific machine noise versus another.
Here is how it works:
- Record many short examples of each sound (at least 20 per class, 1–2 seconds).
- Include background noise as its own class — otherwise, silence will be classified as one of the sounds.
- Train.
- Test with new recordings the model has not heard before.
The model does not see the sound as you hear it — it sees the waveform, often converted into an image (spectrogram) that it processes like an image model.
Interactive
Experiment with three classes (Teachable Machine has an audio mode):
| Class | Number of examples |
|---|---|
| click | 25 |
| snap | 25 |
| background | 25 |
Then deliberately test difficult cases:
- Click further away from the microphone.
- Click with music in the background.
- Have someone else click.
Most models perform significantly worse on all three. Why? All training examples were recorded close up, in a quiet room, by the same person. The model learned the recording situation, not just the sound.
Collect data again with variation and measure again. It is the same lesson as with images — but it feels clearer with audio.
Mastery means
- Collects audio examples and trains a classifier
- Tests with new sounds and explains the errors
Sign in to do the exercises and build your mastery up.
Sources
- Teachable Machine (Google, gratis) — free web service
- Wikipedia — Spektrogram (CC BY-SA 4.0) — CC BY-SA 4.0