Skip to content
AI-grafen
EUniversityGenerative models· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Generative models — an overview

Be able to explain the difference between autoregressive models, VAEs, GANs and diffusion, and what each of them models.

Prerequisites

Intuition

Every generative model tries to learn the data distribution p(x) — how likely different images or texts are — so that new examples can be drawn from it. They differ in how.

FamilyThe ideaGood atWeakness
Autoregressivefactorise p(x) = Πp(xₜ|x₍₌ₜ₎) and predict one piece at a timetext, code; an exact likelihoodslow generation (one token at a time)
VAEencode to a latent distribution, decode backa structured latent space, fastblurry images
GANa generator against a discriminator in a gamesharp images, fast samplingunstable training, mode collapse
Diffusionlearn to remove noise, step by steptop-class images, audio and videomany steps = slow

Language models are autoregressive. Image generators are today nearly always diffusion.

Formal

Autoregressive: pθ(x)=∏tpθ(xt∣x<t)p_\theta(x) = \prod_t p_\theta(x_t\mid x_{<t}), trained with maximum likelihood (cross-entropy). An exact likelihood — which is why perplexity can be computed.

VAE: introduce a latent variable zz and maximise a lower bound (the ELBO): log⁡p(x)≥Eq(z∣x)[log⁡p(x∣z)]−DKL(q(z∣x) ∥ p(z))\log p(x) \ge \mathbb E_{q(z|x)}[\log p(x|z)] - D_{KL}(q(z|x)\,\|\,p(z)). The first term is the reconstruction, the second pulls the latent distribution towards a normal distribution. The blurriness comes from the reconstruction term often being a pixel-wise MSE, which rewards averages.

GAN: a minimax between the generator GG and the discriminator DD: min⁡Gmax⁡DEx[log⁡D(x)]+Ez[log⁡(1−D(G(z)))]\min_G\max_D \mathbb E_x[\log D(x)] + \mathbb E_z[\log(1 - D(G(z)))]. No likelihood — hence no comparable perplexity figures, and hence it is hard to know when the training is going well.

Diffusion: a forward process adds Gaussian noise over TT steps until the image is pure noise; the model learns the reverse process by predicting the noise, with a simple MSE loss. Stable training, but sampling needs many denoising steps (accelerated with DDIM, distillation, consistency models).

Mastery means

  • Tells autoregressive models, VAEs, GANs and diffusion apart
  • States what each family models and where it fits

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences