Skip to content
AI-grafen
FAI engineeringAudio and speech· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Music and audio generation

Be able to explain how generative audio models work and their licensing and copyright questions.

Prerequisites

Intuition

Audio is harder to generate than images for one simple reason: 44 100 samples per second. Three minutes of music is nearly eight million numbers. Generating them one at a time is hopeless.

The solution is the same as for images: compress first, generate in the compressed space.

Two main lines:

ApproachHowExamples
Audio tokensa neural codec (EnCodec, SoundStream) compresses audio to ~50 discrete tokens per second; a transformer generates the tokens; the codec decodes them backMusicGen, AudioLM
Latent diffusionan autoencoder gives a latent space; diffusion generates there; the decoder gives the waveformStable Audio, AudioLDM

Both are the same idea as in image generation, with the same trade-off: harder compression gives faster generation and worse detail.

What is uniquely hard about music: long-range structure. A model can produce eight seconds that sound excellent and still lack what makes three minutes into a song — recurring themes, build-up, resolution.

Formal

Residual vector quantisation (RVQ) is the core of audio codecs. A single codebook is not enough for high quality, so several are layered: the first quantises the signal coarsely, the second quantises the residual, the third the residual of that, and so on. With 8 codebooks of 1 024 entries each and 50 frames per second that is 400 tokens per second — manageable for a transformer, and scalable: fewer layers give a lower bitrate and gradually worse quality.

Evaluating generated audio:

MetricMeasuresWeakness
FAD (Fréchet Audio Distance)the distance between distributions of audio embeddingssays nothing about individual clips
CLAP scorehow well the audio matches the text promptinherited from a model that can itself be wrong
Listening testsquality and musicalityexpensive, but irreplaceable
A memorisation checkthe nearest neighbour in the training datalegally the most important

The legal position — three distinct questions:

  1. The training data. In the EU the TDM exception applies (the DSM Directive art. 4) with the possibility for rights holders to reserve their rights. The music industry has reserved broadly, and several large disputes are ongoing internationally.
  2. Style. A musical style is not protected by copyright. A melody is. The boundary is being tested right now.
  3. The output. Purely machine-generated music probably has no copyright protection in the EU, since human creation is required. That is rarely what the user expects.

Practically for anyone using the tools: read the provider's terms (they differ greatly in what may be used commercially), run a memorisation check if you publish, and be aware that you probably do not own the result on your own — or at all.

Mastery means

  • Explains audio tokens and latent diffusion for audio
  • Evaluates generated audio
  • Reasons about the copyright of the training data and of the output

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences