Skip to content
AI-grafen
FAI engineeringGenerative models· about 90 min· fast-moving, sources checked often· verified 2026-09-21· EN

Latent diffusion and text-to-image

Be able to explain how text-guided image generation works and run a model locally.

Prerequisites

Intuition

Diffusion in a couple of sentences: teach a model to remove noise. Train it by adding noise in steps to real images and letting the model predict the noise. Generate by starting in pure noise and removing a little at a time.

Latent diffusion adds a decisive step: do it in a compressed space instead of in pixels.

image 512×512×3 ──VAE encoder──→ latent 64×64×4 ──diffusion here──→ VAE decoder ──→ image
   786 432 numbers                  16 384 numbers

A factor of 48 fewer numbers to work with. That is what made image generation possible on ordinary graphics cards — Stable Diffusion instead of models that needed a data centre.

Three parts in a text-to-image model:

PartDoes
The text encoder (CLIP or T5)text → vectors
A U-Net or a DiTremoves noise, guided by the text vectors via cross-attention
The VAEbetween pixel space and latent space

Formal

The forward process adds noise according to a schedule:

xt=αˉt x0+1−αˉt ε,ε∼N(0,I)x_t = \sqrt{\bar\alpha_t}\,x_0 + \sqrt{1-\bar\alpha_t}\,\varepsilon, \qquad \varepsilon \sim \mathcal{N}(0, I)

The nice thing is that xtx_t can be computed directly for any tt — no iteration is needed during training.

The training objective is surprisingly simple:

L=Ex0,ε,t[∥ε−εθ(xt,t,c)∥2]\mathcal{L} = \mathbb{E}_{x_0, \varepsilon, t}\left[\left\|\varepsilon - \varepsilon_\theta(x_t, t, c)\right\|^2\right]

The model is to guess which noise was added, given the noisy image, the time step and the conditioning cc (the text embedding). No adversarial training, no instability — that is one of the reasons diffusion beat GANs.

Classifier-free guidance is what makes the text guidance strong. During training the conditioning is dropped at random in about 10 % of cases, so that the model learns both conditional and unconditional generation. At sampling time it extrapolates:

ε~=εθ(xt,∅)+w(εθ(xt,c)−εθ(xt,∅))\tilde\varepsilon = \varepsilon_\theta(x_t, \varnothing) + w\left(\varepsilon_\theta(x_t, c) - \varepsilon_\theta(x_t, \varnothing)\right)

wwEffect
1no guidance
5–8typical — follows the prompt well
> 15oversaturated colours, artefacts, lost variety

The cost is that every step requires two model calls.

Samplers decide how many steps are needed:

SamplerStepsComment
DDPM1 000the original, impractical
DDIM20–50deterministic, reproducible
DPM-Solver++15–25the fastest for good quality
Consistency models, LCM1–4distilled, lower quality

Guidance beyond text:

MethodGuides with
ControlNetan edge image, a depth map, a pose
Inpaintinga mask — generate only in one area
img2imga starting image plus a noise strength
LoRAa learnt style or character
IP-Adaptera reference image for the style

The legal and the ethical belong together with the technology: the copyright of the training data, memorisation of individual images, and labelling synthetic content under Article 50 of the AI Act. Run a memorisation check before you publish.

Code

import torch
from diffusers import StableDiffusionXLPipeline, DPMSolverMultistepScheduler

pipe = StableDiffusionXLPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    torch_dtype=torch.float16, variant="fp16").to("cuda")
pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config)
pipe.enable_attention_slicing()          # less memory, somewhat slower

generator = torch.Generator("cuda").manual_seed(42)      # a seed → reproducible
image = pipe(
    prompt="an educational diagram of a neural network architecture, clean line art, white background",
    negative_prompt="blurry, text, watermark, low resolution",
    num_inference_steps=25,
    guidance_scale=7.0,
    generator=generator,
).images[0]
image.save("diagram.png")

# The guidance scale: measure the effect instead of guessing
for w in (1.0, 3.0, 7.0, 15.0, 25.0):
    b = pipe(prompt="a red cube on a blue table", guidance_scale=w,
             num_inference_steps=25,
             generator=torch.Generator("cuda").manual_seed(0)).images[0]
    b.save(f"cfg_{w}.png")
# w=1: ignores the prompt  ·  w=7: follows it  ·  w=25: oversaturated and broken up

# Steps against quality and time
import time
for steps in (10, 20, 25, 50):
    t0 = time.perf_counter()
    pipe(prompt="a cat", num_inference_steps=steps,
         generator=torch.Generator("cuda").manual_seed(0))
    print(f"  {steps:>2} steps: {time.perf_counter() - t0:.1f} s")
#  ↑ above about 25 steps the improvement is marginal with DPM-Solver++

# The memory requirement
def memory_gb(fp16=True):
    parts = {"U-Net": 2.6, "text encoder": 0.8, "VAE": 0.2}     # SDXL, approximately
    factor = 1.0 if fp16 else 2.0
    return {k: round(v * factor, 2) for k, v in parts.items()} | {
        "total": round(sum(parts.values()) * factor, 2)}
print(memory_gb())    # {'U-Net': 2.6, ..., 'total': 3.6}

# A memorisation check before publishing
def resembles_training_data(image, index, clip_model, threshold=0.95):
    v = clip_model.encode_image(image)
    v = v / v.norm()
    hits = index.search(v, k=5)
    return [t for t in hits if t["similarity"] > threshold]
# A non-empty list → generate again; the image lies too close to a training example.

Mastery means

  • Explains diffusion in a latent space
  • Describes how the text guidance works
  • Chooses the sampler and the number of steps deliberately

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences