Choosing and fine-tuning embedding models
Be able to compare embedding models on your own data and fine-tune with a contrastive loss.
Prerequisites
- DEmbeddings — words as vectorsrequired
- EFine-tuning language modelsrequired
Intuition
The choice of embedding model affects RAG more than the choice of LLM. Four dimensions to compare:
| The dimension | The questions |
|---|---|
| The quality on your data | recall@10 on your own test set — not on MTEB |
| The language | is Swedish in the training data? multilingual or English-centric? |
| The dimension | 384, 768, 1 024, 4 096 — it affects the memory and the speed linearly |
| The context length | 512 tokens is not enough for long chunks |
The MTEB ranking is a starting point, not an answer. Models are optimised against it, and your domain is not in it. Always run your own test set — it takes an hour and gives a different answer surprisingly often.
Formal
Contrastive training is how embedding models are taught: a question and a relevant passage should lie close together, a question and an irrelevant passage far apart. The InfoNCE loss:
(the temperature, often 0.02–0.05) controls how hard the model punishes near-miss negative examples.
Hard negatives are decisive. Random negative examples are too easy — the model learns nothing from telling «how do I change my password» apart from «a recipe for pancakes». Take passages the current model ranks highly but that are wrong. They give the most signal per example.
When does fine-tuning pay off? When you have ≥ 1 000 pairs of (question, the right passage) from your domain — often extractable from logs or generated with an LLM and reviewed. A typical gain: 5–15 percentage points of recall@10 on domain data. With fewer pairs than that, hybrid retrieval and better chunking are cheaper routes to the same gain.
Code
import numpy as np
from sentence_transformers import SentenceTransformer, InputExample, losses
from torch.utils.data import DataLoader
# 1. Compare the candidates on YOUR OWN test set
def compare(models, questions, passages, key, k=10):
for name in models:
m = SentenceTransformer(name)
P = m.encode(passages, normalize_embeddings=True)
Q = m.encode(questions, normalize_embeddings=True)
topk = np.argsort(-(Q @ P.T), axis=1)[:, :k]
recall = np.mean([f in row for row, f in zip(topk, key)])
print(f"{name:45s} recall@{k} {recall:.3f} dim {P.shape[1]}")
compare(["intfloat/multilingual-e5-large", "KBLab/sentence-bert-swedish-cased",
"BAAI/bge-m3"], questions, passages, key)
# 2. Fine-tune with hard negatives
examples = [InputExample(texts=[q, p_pos, p_hard_neg]) for q, p_pos, p_hard_neg in triples]
m = SentenceTransformer("intfloat/multilingual-e5-large")
m.fit(train_objectives=[(DataLoader(examples, batch_size=16, shuffle=True),
losses.MultipleNegativesRankingLoss(m))],
epochs=2, warmup_steps=100)
The trap: e5 and bge models require prefixes ("query: " and "passage: " respectively). If you forget them you lose several percentage points — and it does not show up as an error anywhere.
Mastery means
- Compares embedding models on their own data
- Understands contrastive training and hard negatives
- Knows when fine-tuning pays off
Sign in to do the exercises and build your mastery up.
Sources
- MTEB: Massive Text Embedding Benchmark — free to read
- arXiv — Text Embeddings by Weakly-Supervised Contrastive Pre-training (E5) — arXiv (open access; licence per article)
- Sentence-Transformers (Apache-2.0) — Apache-2.0