Music and audio generation
Be able to explain how generative audio models work and their licensing and copyright questions.
Prerequisites
- FDiffusion modelsrequired
- FSpeech synthesis (TTS)required
Intuition
Audio is harder to generate than images for one simple reason: 44 100 samples per second. Three minutes of music is nearly eight million numbers. Generating them one at a time is hopeless.
The solution is the same as for images: compress first, generate in the compressed space.
Two main lines:
| Approach | How | Examples |
|---|---|---|
| Audio tokens | a neural codec (EnCodec, SoundStream) compresses audio to ~50 discrete tokens per second; a transformer generates the tokens; the codec decodes them back | MusicGen, AudioLM |
| Latent diffusion | an autoencoder gives a latent space; diffusion generates there; the decoder gives the waveform | Stable Audio, AudioLDM |
Both are the same idea as in image generation, with the same trade-off: harder compression gives faster generation and worse detail.
What is uniquely hard about music: long-range structure. A model can produce eight seconds that sound excellent and still lack what makes three minutes into a song — recurring themes, build-up, resolution.
Formal
Residual vector quantisation (RVQ) is the core of audio codecs. A single codebook is not enough for high quality, so several are layered: the first quantises the signal coarsely, the second quantises the residual, the third the residual of that, and so on. With 8 codebooks of 1 024 entries each and 50 frames per second that is 400 tokens per second — manageable for a transformer, and scalable: fewer layers give a lower bitrate and gradually worse quality.
Evaluating generated audio:
| Metric | Measures | Weakness |
|---|---|---|
| FAD (Fréchet Audio Distance) | the distance between distributions of audio embeddings | says nothing about individual clips |
| CLAP score | how well the audio matches the text prompt | inherited from a model that can itself be wrong |
| Listening tests | quality and musicality | expensive, but irreplaceable |
| A memorisation check | the nearest neighbour in the training data | legally the most important |
The legal position — three distinct questions:
- The training data. In the EU the TDM exception applies (the DSM Directive art. 4) with the possibility for rights holders to reserve their rights. The music industry has reserved broadly, and several large disputes are ongoing internationally.
- Style. A musical style is not protected by copyright. A melody is. The boundary is being tested right now.
- The output. Purely machine-generated music probably has no copyright protection in the EU, since human creation is required. That is rarely what the user expects.
Practically for anyone using the tools: read the provider's terms (they differ greatly in what may be used commercially), run a memorisation check if you publish, and be aware that you probably do not own the result on your own — or at all.
Mastery means
- Explains audio tokens and latent diffusion for audio
- Evaluates generated audio
- Reasons about the copyright of the training data and of the output
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Simple and Controllable Music Generation (MusicGen) — arXiv (open access; licence per article)
- arXiv — High Fidelity Neural Audio Compression (EnCodec) — arXiv (open access; licence per article)
- DSM-direktivet (EU) 2019/790, art. 3–4 — EU legal act