DAI developerLab· about 45 min· server sandbox
Lab: k-nearest neighbours from scratch
Implement kNN with Euclidean distance and a majority vote, measure accuracy on sklearn digits and see how k matters.
Theory
No training: for a new point, find the k nearest training points and vote. The distance decides everything — which is why scaling matters.
Sub-tasks
- distances —
distances(x, X)— Euclidean distance from x to every row in X (vectorised). - predict —
predict(Xtr, ytr, X, k)— majority vote among the k nearest (on a tie: the smallest label).
Passes when: accuracy >= 0.95
The starter code
runs in an isolated sandbox on the serverimport numpy as np
def distances(x, X):
# TODO: sqrt(sum((X - x)^2, axis=1))
...
def predict(Xtr, ytr, X, k=3):
# TODO: för varje rad i X: k närmaste index (argsort), majoritetsröst via np.bincount
...
You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.
Try the diagnosticCreate a free accountExpected results
Accuracy ≥ 0.95 on digits with k = 3 (typically ≈ 0.98).
Common mistakes
- Forgets the square root (still works for ranking, but the test requires the real distance).
- Counts the test point itself when Xtr == X.
- Ties are handled inconsistently.