Skip to content
AI-grafen
FAI engineeringRAG and information retrieval· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Choosing and fine-tuning embedding models

Be able to compare embedding models on your own data and fine-tune with a contrastive loss.

Prerequisites

Intuition

The choice of embedding model affects RAG more than the choice of LLM. Four dimensions to compare:

The dimensionThe questions
The quality on your datarecall@10 on your own test set — not on MTEB
The languageis Swedish in the training data? multilingual or English-centric?
The dimension384, 768, 1 024, 4 096 — it affects the memory and the speed linearly
The context length512 tokens is not enough for long chunks

The MTEB ranking is a starting point, not an answer. Models are optimised against it, and your domain is not in it. Always run your own test set — it takes an hour and gives a different answer surprisingly often.

Formal

Contrastive training is how embedding models are taught: a question and a relevant passage should lie close together, a question and an irrelevant passage far apart. The InfoNCE loss:

L=−log⁡exp⁡(sim(q,p+)/τ)exp⁡(sim(q,p+)/τ)+∑p−exp⁡(sim(q,p−)/τ)\mathcal L = -\log\frac{\exp(\text{sim}(q, p^+)/\tau)}{\exp(\text{sim}(q,p^+)/\tau) + \sum_{p^-}\exp(\text{sim}(q,p^-)/\tau)}

τ\tau (the temperature, often 0.02–0.05) controls how hard the model punishes near-miss negative examples.

Hard negatives are decisive. Random negative examples are too easy — the model learns nothing from telling «how do I change my password» apart from «a recipe for pancakes». Take passages the current model ranks highly but that are wrong. They give the most signal per example.

When does fine-tuning pay off? When you have ≥ 1 000 pairs of (question, the right passage) from your domain — often extractable from logs or generated with an LLM and reviewed. A typical gain: 5–15 percentage points of recall@10 on domain data. With fewer pairs than that, hybrid retrieval and better chunking are cheaper routes to the same gain.

Code

import numpy as np
from sentence_transformers import SentenceTransformer, InputExample, losses
from torch.utils.data import DataLoader

# 1. Compare the candidates on YOUR OWN test set
def compare(models, questions, passages, key, k=10):
    for name in models:
        m = SentenceTransformer(name)
        P = m.encode(passages, normalize_embeddings=True)
        Q = m.encode(questions, normalize_embeddings=True)
        topk = np.argsort(-(Q @ P.T), axis=1)[:, :k]
        recall = np.mean([f in row for row, f in zip(topk, key)])
        print(f"{name:45s} recall@{k} {recall:.3f}  dim {P.shape[1]}")

compare(["intfloat/multilingual-e5-large", "KBLab/sentence-bert-swedish-cased",
         "BAAI/bge-m3"], questions, passages, key)

# 2. Fine-tune with hard negatives
examples = [InputExample(texts=[q, p_pos, p_hard_neg]) for q, p_pos, p_hard_neg in triples]
m = SentenceTransformer("intfloat/multilingual-e5-large")
m.fit(train_objectives=[(DataLoader(examples, batch_size=16, shuffle=True),
                         losses.MultipleNegativesRankingLoss(m))],
      epochs=2, warmup_steps=100)

The trap: e5 and bge models require prefixes ("query: " and "passage: " respectively). If you forget them you lose several percentage points — and it does not show up as an error anywhere.

Mastery means

  • Compares embedding models on their own data
  • Understands contrastive training and hard negatives
  • Knows when fine-tuning pays off

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences