Skip to content
AI-grafen
EUniversityMultimodal models· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Multimodal models — the basics

Be able to explain how images, text and sound can meet in a shared vector space.

Prerequisites

Intuition

The idea is simple: train two encoders so that they put things that belong together in the same place.

An image encoder turns an image into a vector. A text encoder turns a text into a vector. Train them together on millions of (image, caption) pairs with the rule: the vector for an image should lie close to the vector for its own caption and far from everybody else's.

Once that is done you can:

  • search for images with text — encode the query, find the nearest image vectors,
  • classify with no training data — encode «a photo of a cat», «a photo of a dog» and see which is closest,
  • connect modalities that were never seen together, if both have been connected to text.

The same recipe works for audio and text (CLAP), video and text, and in principle any two streams where naturally paired data exists.

Formal

The contrastive loss (InfoNCE). In a batch of NN pairs all the vectors are normalised to length 1, and the similarity matrix becomes Sij=ii⋅tj/τS_{ij} = \mathbf{i}_i \cdot \mathbf{t}_j / \tau. The loss is cross-entropy in both directions — the right text for every image and the right image for every text:

L=12[CE(S,diag)+CE(S⊤,diag)]\mathcal{L} = \tfrac{1}{2}\left[\mathrm{CE}(S, \mathrm{diag}) + \mathrm{CE}(S^\top, \mathrm{diag})\right]

The temperature τ\tau is learnt and typically ends up around 0.01. The batch size is decisive: the other N−1N-1 pairs are the negative examples, so a larger batch gives harder and more informative training. CLIP was trained with a batch size of 32 768.

Three limitations that are easy to miss:

  1. A bag of concepts. A shared space captures what is in the image, not how things relate. «The dog chases the cat» and «the cat chases the dog» end up in almost the same place. Compositionality is a known weakness.
  2. No counting. «Three apples» and «five apples» barely differ.
  3. Inherited bias. The space reflects the web data it was trained on, including its stereotypes — and therefore the search results too.

Two architecture families, different purposes:

Two-tower (CLIP)Fused (LLaVA)
Encodersseparate, compared with a dot productimage patches are projected into the language model's token stream
Handlessearch at scale, a pre-encoded indexreasoning and free-form answers about the image
Cost per queryvery lowhigh

They do not compete — in practice CLIP is used to find and a VLM to reason about what was found.

Code

import numpy as np, torch
from transformers import CLIPModel, CLIPProcessor

m = CLIPModel.from_pretrained("openai/clip-vit-base-patch32")
pro = CLIPProcessor.from_pretrained("openai/clip-vit-base-patch32")

def image_vectors(images):
    with torch.no_grad():
        v = m.get_image_features(**pro(images=images, return_tensors="pt"))
    return torch.nn.functional.normalize(v, dim=-1).numpy()

def text_vector(text):
    with torch.no_grad():
        v = m.get_text_features(**pro(text=[text], return_tensors="pt", padding=True))
    return torch.nn.functional.normalize(v, dim=-1).numpy()[0]

INDEX = image_vectors(images)                      # (N, 512), normalised

def search(query, k=5):
    q = text_vector(f"a photo of {query}")          # the prompt template raises the accuracy
    s = INDEX @ q                                   # cosine similarity
    return [(i, float(s[i])) for i in np.argsort(-s)[:k]]

# Classification with no training data
classes = ["a cat", "a dog", "a horse"]
T = np.stack([text_vector(f"a photo of {c}") for c in classes])
print(classes[int(np.argmax(image_vectors([image])[0] @ T.T))])

# The compositionality test — run it on your own data before you trust the space
a, b = text_vector("the dog chases the cat"), text_vector("the cat chases the dog")
print(round(float(a @ b), 3))      # ~0.95 — the model barely sees the difference

The last line is worth running once: it shows more clearly than any explanation why a shared vector space is not an understanding of the scene.

Mastery means

  • Explains contrastive learning across two modalities
  • Searches between modalities with a shared space
  • Knows what a shared space cannot do

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences