Skip to content
AI-grafen
EUniversityRAG and information retrieval· about 60 min· fast-moving, sources checked often· verified 2026-09-20· EN

Vector databases and indexing

Be able to choose an index (HNSW, IVF), understand the recall/latency trade-off, and use Qdrant or pgvector.

Prerequisites

Intuition

Comparing a query vector with a million document vectors exactly takes ~100 ms on a CPU — too slow per query. Approximate nearest neighbour search (ANN) trades a little recall for a lot of speed.

  • HNSW: a graph in several layers where every vector links to its neighbours; the search hops greedily downwards. Fast, high recall, memory-hungry (the graph lives in RAM). Parameters: M (the neighbours), ef (the search width — higher = better recall, slower).
  • IVF: cluster the vectors; search only in the nprobe nearest clusters. Less memory, slightly lower recall.
  • Quantisation (PQ/scalar) shrinks the vectors 4–32× with little loss of quality.

Metadata filters (language, licence, date) should be applied in the index, not afterwards — otherwise the top k can be empty after filtering. Always measure recall@k against brute force on a sample.

Code

from qdrant_client import QdrantClient, models as qm
import numpy as np

q = QdrantClient("http://localhost:6333")
q.recreate_collection("docs", vectors_config=qm.VectorParams(size=768, distance=qm.Distance.COSINE),
                      hnsw_config=qm.HnswConfigDiff(m=16, ef_construct=128))
q.upsert("docs", points=[qm.PointStruct(id=i, vector=v.tolist(), payload={"lang": "sv", "license": "CC-BY"}) for i, v in enumerate(vecs)])

hits = q.search("docs", query_vector=qv.tolist(), limit=10, search_params=qm.SearchParams(hnsw_ef=64),
                query_filter=qm.Filter(must=[qm.FieldCondition(key="license", match=qm.MatchValue(value="CC-BY"))]))

# recall against the exact search
exact = np.argsort(-(vecs @ qv))[:10]
print(len(set(h.id for h in hits) & set(exact)) / 10)

pgvector: CREATE INDEX ON chunk USING hnsw (emb vector_cosine_ops); SELECT … ORDER BY emb <=> $1 LIMIT 10 with WHERE license = 'CC-BY' in the same query.

Mastery means

  • Explains HNSW and IVF and the recall/latency trade-off
  • Uses Qdrant or pgvector with metadata filters
  • Measures recall against an exact search

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences