DAI developerLab· about 45 min· server sandbox
Lab: k-means from scratch
Implement k-means (assign, update, repeat), show that the inertia never increases, and find clusters in synthetic data.
Theory
Repeat: assign every point to the nearest centroid; move every centroid to the mean of its points. The inertia (sum of squared distances) decreases monotonically.
Sub-tasks
- assign —
assign(X, C)→ index of the nearest centroid per row. - update —
update(X, labels, k)→ new centroids (an empty cluster falls back via NaN handling: use the mean of X). - kmeans —
kmeans(X, k, steps, seed)→ (C, labels, inertia_history).
Passes when: inertia <= 250
The starter code
runs in an isolated sandbox on the serverimport numpy as np
def assign(X, C):
# TODO: avstånd (n, k) → argmin per rad
...
def update(X, labels, k):
# TODO: medelvärde per kluster; tomt kluster → X.mean(axis=0)
...
def inertia(X, C, labels):
return float(((X - C[labels]) ** 2).sum())
def kmeans(X, k=3, steps=20, seed=0):
rng = np.random.default_rng(seed)
C = X[rng.choice(len(X), k, replace=False)]
hist = []
for _ in range(steps):
# TODO: assign → hist.append(inertia) → update
...
return C, assign(X, C), hist
You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.
Try the diagnosticCreate a free accountExpected results
The inertia decreases monotonically; on three clear blobs the right clusters are found (inertia < 250).
Common mistakes
- Swaps the axis in the distance computation.
- An empty cluster gives a NaN centroid.
- Initialises centroids outside the data.