Skip to content
AI-grafen
FAI engineeringMultimodal models· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Multimodal RAG

Be able to index and retrieve images and tables alongside text.

Prerequisites

Intuition

Ordinary RAG retrieves passages of text. But the answer to «what does the trend look like since 2020?» is often in a chart, and the answer to «what does the subscription cost?» in a table.

Three strategies for non-text:

StrategyWhat gets indexedStrengthWeakness
Describe and index the texta VLM-generated description of every imageworks with your existing text searchthe description becomes a filter — what it does not mention cannot be found
A shared vector spaceimage vectors in the same space as the text queryfinds things visually with no intermediateblunt for text inside the image and for details
The page image directlythe whole page as an image (ColPali style)no loss of informationa larger index, more expensive

In practice a hybrid nearly always wins: index both the description and the image vector, retrieve from both, and merge.

The golden rule: always keep the original. Retrieve on the description, but send the image to the model that has to answer. Otherwise it is answering from a summary of a summary.

Formal

Reciprocal rank fusion (RRF) is the standard way of merging rankings from different indexes:

score(d)=∑r∈R1k+rankr(d)\mathrm{score}(d) = \sum_{r \in R} \frac{1}{k + \mathrm{rank}_r(d)}

with k≈60k \approx 60. It has two properties that make it hard to beat in practice: it needs no comparable scores between the indexes (only the order), and it rewards consensus — a document that comes second in two indexes beats one that comes first in one and fourth in another.

Evaluate the retrieval and the generation separately. Otherwise you do not know what to fix:

StepMetricThe question
Retrievalrecall@k per modalitywas the right material among the hits?
RankingMRR, nDCGwas it high enough up?
Generationgroundedness against the retrieved materialwas what was retrieved actually used?
End to endanswer accuracywas the answer right?

Measure the recall per modality separately. The most common hidden fault in multimodal RAG is a text recall of 0.9 and an image recall of 0.3, while the overall figure looks acceptable — and every question that needs a chart fails silently.

Four practical decisions:

  1. Chunking that respects the layout. Do not cut in the middle of a table; keep the caption with the image.
  2. Keep the relationship. An image should know which text surrounds it; a table which heading it belongs to.
  3. Pass the original on. The description is a way in, not the material.
  4. Do not mix spaces. Two different embedding models do not give comparable distances — keep the indexes apart and merge on rank.

Code

from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams, PointStruct

cl = QdrantClient(url="http://qdrant:6333")
for name, dim in (("text", 768), ("image", 512)):
    cl.recreate_collection(name, vectors_config=VectorParams(size=dim, distance=Distance.COSINE))

def index_page(page):
    # 1. the text as usual
    for i, passage in enumerate(page["passages"]):
        cl.upsert("text", [PointStruct(id=f"{page['id']}-t{i}", vector=text_emb(passage),
                                       payload={"page": page["id"], "type": "text", "content": passage})])
    # 2. images: BOTH a description in the text index AND an image vector in the image index
    for j, image in enumerate(page["images"]):
        desc = vlm_describe(image, page["captions"][j])
        cl.upsert("text", [PointStruct(id=f"{page['id']}-it{j}", vector=text_emb(desc),
                                       payload={"page": page["id"], "type": "image",
                                                "image_path": image.path, "description": desc})])
        cl.upsert("image", [PointStruct(id=f"{page['id']}-i{j}", vector=clip_image(image),
                                        payload={"page": page["id"], "type": "image",
                                                 "image_path": image.path})])

def rrf(rankings, k=60, top=8):
    scores = {}
    for r in rankings:
        for pos, d in enumerate(r):
            scores[d] = scores.get(d, 0.0) + 1.0 / (k + pos + 1)
    return sorted(scores, key=lambda d: (-scores[d], d))[:top]

def retrieve(query, k=8):
    t = [p.id for p in cl.search("text", text_emb(query), limit=k)]
    i = [p.id for p in cl.search("image", clip_text(query), limit=k)]
    return rrf([t, i], top=k)

# Recall per modality — what reveals a silent image failure
def recall_per_modality(truth, k=8):
    out = {}
    for kind in ("text", "image", "table"):
        f = [x for x in truth if x["type"] == kind]
        out[kind] = round(sum(x["correct_id"] in retrieve(x["query"], k) for x in f) / max(len(f), 1), 3)
    return out

print(recall_per_modality(truth))
# {'text': 0.91, 'image': 0.34, 'table': 0.62}   ← the overall figure would have hidden this

Mastery means

  • Indexes several modalities for retrieval
  • Merges rankings from different indexes
  • Evaluates the retrieval and the answer separately

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences