Generative models — an overview
Be able to explain the difference between autoregressive models, VAEs, GANs and diffusion, and what each of them models.
Prerequisites
Intuition
Every generative model tries to learn the data distribution p(x) — how likely different images or texts are — so that new examples can be drawn from it. They differ in how.
| Family | The idea | Good at | Weakness |
|---|---|---|---|
| Autoregressive | factorise p(x) = Πp(xₜ|x₍₌ₜ₎) and predict one piece at a time | text, code; an exact likelihood | slow generation (one token at a time) |
| VAE | encode to a latent distribution, decode back | a structured latent space, fast | blurry images |
| GAN | a generator against a discriminator in a game | sharp images, fast sampling | unstable training, mode collapse |
| Diffusion | learn to remove noise, step by step | top-class images, audio and video | many steps = slow |
Language models are autoregressive. Image generators are today nearly always diffusion.
Formal
Autoregressive: , trained with maximum likelihood (cross-entropy). An exact likelihood — which is why perplexity can be computed.
VAE: introduce a latent variable and maximise a lower bound (the ELBO): . The first term is the reconstruction, the second pulls the latent distribution towards a normal distribution. The blurriness comes from the reconstruction term often being a pixel-wise MSE, which rewards averages.
GAN: a minimax between the generator and the discriminator : . No likelihood — hence no comparable perplexity figures, and hence it is hard to know when the training is going well.
Diffusion: a forward process adds Gaussian noise over steps until the image is pure noise; the model learns the reverse process by predicting the noise, with a simple MSE loss. Stable training, but sampling needs many denoising steps (accelerated with DDIM, distillation, consistency models).
Mastery means
- Tells autoregressive models, VAEs, GANs and diffusion apart
- States what each family models and where it fits
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Denoising Diffusion Probabilistic Models — arXiv (open access; licence per article)
- arXiv — Auto-Encoding Variational Bayes — arXiv (open access; licence per article)
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0