Skip to content
AI-grafen
FAI engineeringComputer vision· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

CLIP and contrastive image–text learning

Be able to explain the contrastive loss and use CLIP for zero-shot classification.

Prerequisites

Intuition

CLIP was trained on 400 million (image, text) pairs from the web with a single, surprisingly simple task: match the right image with the right caption within the batch.

Take 32 768 pairs. Encode all the images and all the texts. Build the similarity matrix — 32 768 × 32 768. The right answers sit on the diagonal. Train with cross-entropy over the rows and over the columns.

It sounds like a curiosity. But because captions on the web describe everything — objects, styles, places, moods, sketches, memes — the model learns a space where almost any concept at all has a direction.

And that gives you something that was not possible before: zero-shot classification. If you want to tell 200 bird species apart, you write 200 sentences and encode them. No training, no labelled images.

Formal

The symmetric loss. With image vectors II and text vectors TT, both L2-normalised, and a learnt temperature τ\tau:

S=IT⊤τ,L=12[CE(S,y)+CE(S⊤,y)],y=(0,1,…,N−1)S = \frac{I T^\top}{\tau}, \qquad \mathcal{L} = \tfrac{1}{2}\left[\mathrm{CE}(S, y) + \mathrm{CE}(S^\top, y)\right], \quad y = (0, 1, \dots, N-1)

Three design choices that matter:

  1. The normalisation turns the dot product into a cosine similarity, bounded to [−1,1][-1, 1]. Without it the logits explode.
  2. The temperature is learnt (as log⁡1/τ\log 1/\tau, clipped at 100) and ends up around 0.01 — a sharp distribution.
  3. The batch is the negative examples. Hence the large batch size; distributed training gathers vectors from every GPU before the matrix is built.

Prompt engineering matters unexpectedly much. «cat» is worse than «a photo of a cat», which is worse than an average over some eighty templates («a blurry photo of a {}», «a sketch of a {}» …). The difference is several percentage points — the training data's captions are whole sentences, not isolated words.

The weaknesses, and how you measure them on your own data:

WeaknessThe test
Compositionalitythe similarity between «A chases B» and «B chases A»
Counting«three X» against «five X» on images with a known count
Text inside imagesan image of the word «apple» is often classified as an apple
Fine-grainednessbird species, dog breeds, components
Biasaccuracy broken down by group
Swedishthe same classes in Swedish against English

The last row is decisive for AI-grafen: the original CLIP was trained on predominantly English captions and performs worse on Swedish prompts. You either translate the class names into English or use a multilingual variant — but measure the difference before you choose.

Code

import torch, torch.nn.functional as F

def clip_loss(image_v, text_v, logit_scale):
    """image_v, text_v: (N, d). logit_scale: exp(the learnt parameter), clipped at 100."""
    I = F.normalize(image_v, dim=-1)
    T = F.normalize(text_v, dim=-1)
    S = logit_scale * I @ T.t()                       # (N, N)
    y = torch.arange(len(S), device=S.device)
    return (F.cross_entropy(S, y) + F.cross_entropy(S.t(), y)) / 2

# A check: a perfect model gives near zero, a random one gives log(N)
N, d = 8, 16
perfect = torch.randn(N, d)
print(round(float(clip_loss(perfect, perfect.clone(), 100.0)), 4))           # 0.0
print(round(float(clip_loss(torch.randn(N, d), torch.randn(N, d), 1.0)), 3),
      round(float(torch.tensor(N).float().log()), 3))                        # ~2.0  2.079

# Zero-shot with a prompt ensemble
TEMPLATES = ["a photo of a {}", "a picture of a {}", "a close-up of a {}",
             "a blurry photo of a {}", "a drawing of a {}"]

def class_vectors(classes, encode_text):
    out = []
    for c in classes:
        v = torch.stack([encode_text(t.format(c)) for t in TEMPLATES])
        v = F.normalize(v, dim=-1).mean(0)
        out.append(F.normalize(v, dim=0))
    return torch.stack(out)                            # (K, d)

# The compositionality test — run it before you trust the space
a = encode_text("a dog chasing a cat")
b = encode_text("a cat chasing a dog")
print(round(float(F.cosine_similarity(a[None], b[None])), 3))    # ~0.95

The check on lines 12–14 is worth keeping in the test suite: a correct implementation should give zero on identical vectors and log⁡N\log N on random ones. Errors in the normalisation or the transposition show up there immediately.

Mastery means

  • Implements and explains the symmetric contrastive loss
  • Uses CLIP for zero-shot classification
  • Knows CLIP's weaknesses and how they are measured

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences