EUniversityRAG and information retrieval· about 60 min· fast-moving, sources checked often· verified 2026-09-20· EN
Vector databases and indexing
Be able to choose an index (HNSW, IVF), understand the recall/latency trade-off, and use Qdrant or pgvector.
Prerequisites
- DRetrieval — finding the right textrequired
Intuition
Comparing a query vector with a million document vectors exactly takes ~100 ms on a CPU — too slow per query. Approximate nearest neighbour search (ANN) trades a little recall for a lot of speed.
- HNSW: a graph in several layers where every vector links to its neighbours; the search hops greedily downwards. Fast, high recall, memory-hungry (the graph lives in RAM). Parameters:
M(the neighbours),ef(the search width — higher = better recall, slower). - IVF: cluster the vectors; search only in the
nprobenearest clusters. Less memory, slightly lower recall. - Quantisation (PQ/scalar) shrinks the vectors 4–32× with little loss of quality.
Metadata filters (language, licence, date) should be applied in the index, not afterwards — otherwise the top k can be empty after filtering. Always measure recall@k against brute force on a sample.
Code
from qdrant_client import QdrantClient, models as qm
import numpy as np
q = QdrantClient("http://localhost:6333")
q.recreate_collection("docs", vectors_config=qm.VectorParams(size=768, distance=qm.Distance.COSINE),
hnsw_config=qm.HnswConfigDiff(m=16, ef_construct=128))
q.upsert("docs", points=[qm.PointStruct(id=i, vector=v.tolist(), payload={"lang": "sv", "license": "CC-BY"}) for i, v in enumerate(vecs)])
hits = q.search("docs", query_vector=qv.tolist(), limit=10, search_params=qm.SearchParams(hnsw_ef=64),
query_filter=qm.Filter(must=[qm.FieldCondition(key="license", match=qm.MatchValue(value="CC-BY"))]))
# recall against the exact search
exact = np.argsort(-(vecs @ qv))[:10]
print(len(set(h.id for h in hits) & set(exact)) / 10)
pgvector: CREATE INDEX ON chunk USING hnsw (emb vector_cosine_ops); SELECT … ORDER BY emb <=> $1 LIMIT 10 with WHERE license = 'CC-BY' in the same query.
Mastery means
- Explains HNSW and IVF and the recall/latency trade-off
- Uses Qdrant or pgvector with metadata filters
- Measures recall against an exact search
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Efficient and robust approximate nearest neighbor search using HNSW graphs — arXiv (open access; licence per article)
- Qdrant — dokumentation (Apache-2.0) — Apache-2.0