CLIP and contrastive image–text learning
Be able to explain the contrastive loss and use CLIP for zero-shot classification.
Prerequisites
- DEmbeddings — words as vectorsrequired
- EVision Transformer (ViT)required
Intuition
CLIP was trained on 400 million (image, text) pairs from the web with a single, surprisingly simple task: match the right image with the right caption within the batch.
Take 32 768 pairs. Encode all the images and all the texts. Build the similarity matrix — 32 768 × 32 768. The right answers sit on the diagonal. Train with cross-entropy over the rows and over the columns.
It sounds like a curiosity. But because captions on the web describe everything — objects, styles, places, moods, sketches, memes — the model learns a space where almost any concept at all has a direction.
And that gives you something that was not possible before: zero-shot classification. If you want to tell 200 bird species apart, you write 200 sentences and encode them. No training, no labelled images.
Formal
The symmetric loss. With image vectors and text vectors , both L2-normalised, and a learnt temperature :
Three design choices that matter:
- The normalisation turns the dot product into a cosine similarity, bounded to . Without it the logits explode.
- The temperature is learnt (as , clipped at 100) and ends up around 0.01 — a sharp distribution.
- The batch is the negative examples. Hence the large batch size; distributed training gathers vectors from every GPU before the matrix is built.
Prompt engineering matters unexpectedly much. «cat» is worse than «a photo of a cat», which is worse than an average over some eighty templates («a blurry photo of a {}», «a sketch of a {}» …). The difference is several percentage points — the training data's captions are whole sentences, not isolated words.
The weaknesses, and how you measure them on your own data:
| Weakness | The test |
|---|---|
| Compositionality | the similarity between «A chases B» and «B chases A» |
| Counting | «three X» against «five X» on images with a known count |
| Text inside images | an image of the word «apple» is often classified as an apple |
| Fine-grainedness | bird species, dog breeds, components |
| Bias | accuracy broken down by group |
| Swedish | the same classes in Swedish against English |
The last row is decisive for AI-grafen: the original CLIP was trained on predominantly English captions and performs worse on Swedish prompts. You either translate the class names into English or use a multilingual variant — but measure the difference before you choose.
Code
import torch, torch.nn.functional as F
def clip_loss(image_v, text_v, logit_scale):
"""image_v, text_v: (N, d). logit_scale: exp(the learnt parameter), clipped at 100."""
I = F.normalize(image_v, dim=-1)
T = F.normalize(text_v, dim=-1)
S = logit_scale * I @ T.t() # (N, N)
y = torch.arange(len(S), device=S.device)
return (F.cross_entropy(S, y) + F.cross_entropy(S.t(), y)) / 2
# A check: a perfect model gives near zero, a random one gives log(N)
N, d = 8, 16
perfect = torch.randn(N, d)
print(round(float(clip_loss(perfect, perfect.clone(), 100.0)), 4)) # 0.0
print(round(float(clip_loss(torch.randn(N, d), torch.randn(N, d), 1.0)), 3),
round(float(torch.tensor(N).float().log()), 3)) # ~2.0 2.079
# Zero-shot with a prompt ensemble
TEMPLATES = ["a photo of a {}", "a picture of a {}", "a close-up of a {}",
"a blurry photo of a {}", "a drawing of a {}"]
def class_vectors(classes, encode_text):
out = []
for c in classes:
v = torch.stack([encode_text(t.format(c)) for t in TEMPLATES])
v = F.normalize(v, dim=-1).mean(0)
out.append(F.normalize(v, dim=0))
return torch.stack(out) # (K, d)
# The compositionality test — run it before you trust the space
a = encode_text("a dog chasing a cat")
b = encode_text("a cat chasing a dog")
print(round(float(F.cosine_similarity(a[None], b[None])), 3)) # ~0.95
The check on lines 12–14 is worth keeping in the test suite: a correct implementation should give zero on identical vectors and on random ones. Errors in the normalisation or the transposition show up there immediately.
Mastery means
- Implements and explains the symmetric contrastive loss
- Uses CLIP for zero-shot classification
- Knows CLIP's weaknesses and how they are measured
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Learning Transferable Visual Models From Natural Language Supervision (CLIP) — arXiv (open access; licence per article)
- arXiv — When and why vision-language models behave like bags-of-words — arXiv (open access; licence per article)
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0