Multimodal RAG
Be able to index and retrieve images and tables alongside text.
Prerequisites
- EMultimodal models — the basicsrequired
- ERAG — retrieval-augmented generationrequired
Intuition
Ordinary RAG retrieves passages of text. But the answer to «what does the trend look like since 2020?» is often in a chart, and the answer to «what does the subscription cost?» in a table.
Three strategies for non-text:
| Strategy | What gets indexed | Strength | Weakness |
|---|---|---|---|
| Describe and index the text | a VLM-generated description of every image | works with your existing text search | the description becomes a filter — what it does not mention cannot be found |
| A shared vector space | image vectors in the same space as the text query | finds things visually with no intermediate | blunt for text inside the image and for details |
| The page image directly | the whole page as an image (ColPali style) | no loss of information | a larger index, more expensive |
In practice a hybrid nearly always wins: index both the description and the image vector, retrieve from both, and merge.
The golden rule: always keep the original. Retrieve on the description, but send the image to the model that has to answer. Otherwise it is answering from a summary of a summary.
Formal
Reciprocal rank fusion (RRF) is the standard way of merging rankings from different indexes:
with . It has two properties that make it hard to beat in practice: it needs no comparable scores between the indexes (only the order), and it rewards consensus — a document that comes second in two indexes beats one that comes first in one and fourth in another.
Evaluate the retrieval and the generation separately. Otherwise you do not know what to fix:
| Step | Metric | The question |
|---|---|---|
| Retrieval | recall@k per modality | was the right material among the hits? |
| Ranking | MRR, nDCG | was it high enough up? |
| Generation | groundedness against the retrieved material | was what was retrieved actually used? |
| End to end | answer accuracy | was the answer right? |
Measure the recall per modality separately. The most common hidden fault in multimodal RAG is a text recall of 0.9 and an image recall of 0.3, while the overall figure looks acceptable — and every question that needs a chart fails silently.
Four practical decisions:
- Chunking that respects the layout. Do not cut in the middle of a table; keep the caption with the image.
- Keep the relationship. An image should know which text surrounds it; a table which heading it belongs to.
- Pass the original on. The description is a way in, not the material.
- Do not mix spaces. Two different embedding models do not give comparable distances — keep the indexes apart and merge on rank.
Code
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams, PointStruct
cl = QdrantClient(url="http://qdrant:6333")
for name, dim in (("text", 768), ("image", 512)):
cl.recreate_collection(name, vectors_config=VectorParams(size=dim, distance=Distance.COSINE))
def index_page(page):
# 1. the text as usual
for i, passage in enumerate(page["passages"]):
cl.upsert("text", [PointStruct(id=f"{page['id']}-t{i}", vector=text_emb(passage),
payload={"page": page["id"], "type": "text", "content": passage})])
# 2. images: BOTH a description in the text index AND an image vector in the image index
for j, image in enumerate(page["images"]):
desc = vlm_describe(image, page["captions"][j])
cl.upsert("text", [PointStruct(id=f"{page['id']}-it{j}", vector=text_emb(desc),
payload={"page": page["id"], "type": "image",
"image_path": image.path, "description": desc})])
cl.upsert("image", [PointStruct(id=f"{page['id']}-i{j}", vector=clip_image(image),
payload={"page": page["id"], "type": "image",
"image_path": image.path})])
def rrf(rankings, k=60, top=8):
scores = {}
for r in rankings:
for pos, d in enumerate(r):
scores[d] = scores.get(d, 0.0) + 1.0 / (k + pos + 1)
return sorted(scores, key=lambda d: (-scores[d], d))[:top]
def retrieve(query, k=8):
t = [p.id for p in cl.search("text", text_emb(query), limit=k)]
i = [p.id for p in cl.search("image", clip_text(query), limit=k)]
return rrf([t, i], top=k)
# Recall per modality — what reveals a silent image failure
def recall_per_modality(truth, k=8):
out = {}
for kind in ("text", "image", "table"):
f = [x for x in truth if x["type"] == kind]
out[kind] = round(sum(x["correct_id"] in retrieve(x["query"], k) for x in f) / max(len(f), 1), 3)
return out
print(recall_per_modality(truth))
# {'text': 0.91, 'image': 0.34, 'table': 0.62} ← the overall figure would have hidden this
Mastery means
- Indexes several modalities for retrieval
- Merges rankings from different indexes
- Evaluates the retrieval and the answer separately
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — ColPali: Efficient Document Retrieval with Vision Language Models — arXiv (open access; licence per article)
- Cormack m.fl. — Reciprocal Rank Fusion (SIGIR 2009) — abstract free; author copies available
- Qdrant — dokumentation (Apache-2.0) — Apache-2.0