Latent diffusion and text-to-image
Be able to explain how text-guided image generation works and run a model locally.
Prerequisites
- FCLIP and contrastive image–text learningrequired
- FDiffusion modelsrequired
Intuition
Diffusion in a couple of sentences: teach a model to remove noise. Train it by adding noise in steps to real images and letting the model predict the noise. Generate by starting in pure noise and removing a little at a time.
Latent diffusion adds a decisive step: do it in a compressed space instead of in pixels.
image 512×512×3 ──VAE encoder──→ latent 64×64×4 ──diffusion here──→ VAE decoder ──→ image
786 432 numbers 16 384 numbers
A factor of 48 fewer numbers to work with. That is what made image generation possible on ordinary graphics cards — Stable Diffusion instead of models that needed a data centre.
Three parts in a text-to-image model:
| Part | Does |
|---|---|
| The text encoder (CLIP or T5) | text → vectors |
| A U-Net or a DiT | removes noise, guided by the text vectors via cross-attention |
| The VAE | between pixel space and latent space |
Formal
The forward process adds noise according to a schedule:
The nice thing is that can be computed directly for any — no iteration is needed during training.
The training objective is surprisingly simple:
The model is to guess which noise was added, given the noisy image, the time step and the conditioning (the text embedding). No adversarial training, no instability — that is one of the reasons diffusion beat GANs.
Classifier-free guidance is what makes the text guidance strong. During training the conditioning is dropped at random in about 10 % of cases, so that the model learns both conditional and unconditional generation. At sampling time it extrapolates:
| Effect | |
|---|---|
| 1 | no guidance |
| 5–8 | typical — follows the prompt well |
| > 15 | oversaturated colours, artefacts, lost variety |
The cost is that every step requires two model calls.
Samplers decide how many steps are needed:
| Sampler | Steps | Comment |
|---|---|---|
| DDPM | 1 000 | the original, impractical |
| DDIM | 20–50 | deterministic, reproducible |
| DPM-Solver++ | 15–25 | the fastest for good quality |
| Consistency models, LCM | 1–4 | distilled, lower quality |
Guidance beyond text:
| Method | Guides with |
|---|---|
| ControlNet | an edge image, a depth map, a pose |
| Inpainting | a mask — generate only in one area |
| img2img | a starting image plus a noise strength |
| LoRA | a learnt style or character |
| IP-Adapter | a reference image for the style |
The legal and the ethical belong together with the technology: the copyright of the training data, memorisation of individual images, and labelling synthetic content under Article 50 of the AI Act. Run a memorisation check before you publish.
Code
import torch
from diffusers import StableDiffusionXLPipeline, DPMSolverMultistepScheduler
pipe = StableDiffusionXLPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
torch_dtype=torch.float16, variant="fp16").to("cuda")
pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config)
pipe.enable_attention_slicing() # less memory, somewhat slower
generator = torch.Generator("cuda").manual_seed(42) # a seed → reproducible
image = pipe(
prompt="an educational diagram of a neural network architecture, clean line art, white background",
negative_prompt="blurry, text, watermark, low resolution",
num_inference_steps=25,
guidance_scale=7.0,
generator=generator,
).images[0]
image.save("diagram.png")
# The guidance scale: measure the effect instead of guessing
for w in (1.0, 3.0, 7.0, 15.0, 25.0):
b = pipe(prompt="a red cube on a blue table", guidance_scale=w,
num_inference_steps=25,
generator=torch.Generator("cuda").manual_seed(0)).images[0]
b.save(f"cfg_{w}.png")
# w=1: ignores the prompt · w=7: follows it · w=25: oversaturated and broken up
# Steps against quality and time
import time
for steps in (10, 20, 25, 50):
t0 = time.perf_counter()
pipe(prompt="a cat", num_inference_steps=steps,
generator=torch.Generator("cuda").manual_seed(0))
print(f" {steps:>2} steps: {time.perf_counter() - t0:.1f} s")
# ↑ above about 25 steps the improvement is marginal with DPM-Solver++
# The memory requirement
def memory_gb(fp16=True):
parts = {"U-Net": 2.6, "text encoder": 0.8, "VAE": 0.2} # SDXL, approximately
factor = 1.0 if fp16 else 2.0
return {k: round(v * factor, 2) for k, v in parts.items()} | {
"total": round(sum(parts.values()) * factor, 2)}
print(memory_gb()) # {'U-Net': 2.6, ..., 'total': 3.6}
# A memorisation check before publishing
def resembles_training_data(image, index, clip_model, threshold=0.95):
v = clip_model.encode_image(image)
v = v / v.norm()
hits = index.search(v, k=5)
return [t for t in hits if t["similarity"] > threshold]
# A non-empty list → generate again; the image lies too close to a training example.
Mastery means
- Explains diffusion in a latent space
- Describes how the text guidance works
- Chooses the sampler and the number of steps deliberately
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — High-Resolution Image Synthesis with Latent Diffusion Models — arXiv (open access; licence per article)
- arXiv — Classifier-Free Diffusion Guidance — arXiv (open access; licence per article)
- diffusers — dokumentation (Apache-2.0) — Apache-2.0