DAI developerLab· about 45 min· server sandbox
Lab: semantic search with embeddings
Build a small search engine: normalise embeddings, compute cosine similarity against all documents with one matrix product, take the top k, and measure recall.
Teaches: Embeddings — words as vectors
Requires: Vectors
Theory
Normalise all vectors to length 1 — then cosine similarity is just a matrix product Q·Dᵀ. Top k = the k largest per row.
Sub-tasks
- normalize —
normalize(M)divides every row by its norm. - topk —
topk(q, D, k)returns the indices of the k most similar documents (descending). - recall_at_k —
recall_at_k(Q, D, truth, k)— fraction of queries where the right document is among the top k.
Passes when: recall >= 0.9
The starter code
runs in an isolated sandbox on the serverimport numpy as np
def normalize(M):
# TODO: M / norm per rad (undvik division med noll)
...
def topk(q, D, k=3):
# TODO: likheter = normalize(D) @ normalize(q); index till k största, fallande
...
def recall_at_k(Q, D, truth, k=3):
# TODO: andel i där truth[i] finns i topk(Q[i], D, k)
...
You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.
Try the diagnosticCreate a free accountExpected results
recall@3 ≥ 0.9 on the synthetic dataset (queries = documents + noise).
Common mistakes
- Normalises columns instead of rows.
argsortis ascending — reverse the order.- Division by zero for a zero vector.