Skip to content
AI-grafen
DAI developerLab· about 45 min· server sandbox

Lab: semantic search with embeddings

Build a small search engine: normalise embeddings, compute cosine similarity against all documents with one matrix product, take the top k, and measure recall.

Theory

Normalise all vectors to length 1 — then cosine similarity is just a matrix product Q·Dᵀ. Top k = the k largest per row.

Sub-tasks

  1. normalize — normalize(M) divides every row by its norm.
  2. topk — topk(q, D, k) returns the indices of the k most similar documents (descending).
  3. recall_at_k — recall_at_k(Q, D, truth, k) — fraction of queries where the right document is among the top k.

Passes when: recall >= 0.9

The starter code

runs in an isolated sandbox on the server
import numpy as np


def normalize(M):
    # TODO: M / norm per rad (undvik division med noll)
    ...


def topk(q, D, k=3):
    # TODO: likheter = normalize(D) @ normalize(q); index till k största, fallande
    ...


def recall_at_k(Q, D, truth, k=3):
    # TODO: andel i där truth[i] finns i topk(Q[i], D, k)
    ...

You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.

Try the diagnosticCreate a free account

Expected results

recall@3 ≥ 0.9 on the synthetic dataset (queries = documents + noise).

Common mistakes

  • Normalises columns instead of rows.
  • argsort is ascending — reverse the order.
  • Division by zero for a zero vector.