Skip to content
AI-grafen
DAI developerLab· about 45 min· server sandbox

Lab: k-nearest neighbours from scratch

Implement kNN with Euclidean distance and a majority vote, measure accuracy on sklearn digits and see how k matters.

Theory

No training: for a new point, find the k nearest training points and vote. The distance decides everything — which is why scaling matters.

Sub-tasks

  1. distances — distances(x, X) — Euclidean distance from x to every row in X (vectorised).
  2. predict — predict(Xtr, ytr, X, k) — majority vote among the k nearest (on a tie: the smallest label).

Passes when: accuracy >= 0.95

The starter code

runs in an isolated sandbox on the server
import numpy as np


def distances(x, X):
    # TODO: sqrt(sum((X - x)^2, axis=1))
    ...


def predict(Xtr, ytr, X, k=3):
    # TODO: för varje rad i X: k närmaste index (argsort), majoritetsröst via np.bincount
    ...

You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.

Try the diagnosticCreate a free account

Expected results

Accuracy ≥ 0.95 on digits with k = 3 (typically ≈ 0.98).

Common mistakes

  • Forgets the square root (still works for ranking, but the test requires the real distance).
  • Counts the test point itself when Xtr == X.
  • Ties are handled inconsistently.