Clustering: k-means and hierarchical
Be able to cluster data without labels and judge the number of clusters.
Prerequisites
Intuition
Clustering is learning without labels: finding groups in the data that nobody has told you about.
k-means in four steps:
- Place k centre points at random.
- Assign every data point to the nearest centre.
- Move each centre to the mean of its points.
- Repeat 2–3 until nothing changes.
The problem: k has to be chosen in advance, and the algorithm always finds k clusters — even if the data has none.
Hierarchical clustering instead builds a tree: merge the two closest groups, repeat. Then you can cut the tree at any level afterwards.
Code
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
import numpy as np
inertias, silhouettes = [], []
for k in range(2, 9):
km = KMeans(n_clusters=k, n_init=10, random_state=0).fit(X)
inertias.append(km.inertia_) # the sum of squared distances to the centre
silhouettes.append(silhouette_score(X, km.labels_))
for k, i, s in zip(range(2, 9), inertias, silhouettes):
print(k, round(i, 1), round(s, 3))
# 2 520.3 0.51
# 3 210.4 0.68 ← the elbow and the highest silhouette
# 4 195.1 0.42
The elbow method: plot the inertia against k and look for where the curve bends. Silhouette (−1 to 1) measures how well each point fits in its cluster compared with the nearest other one — the highest value is often a good k.
When k-means is the wrong choice: elongated or ring-shaped clusters (it assumes round ones of roughly equal size), differing density, or a lot of noise. Then DBSCAN or hierarchical clustering fits better.
Mastery means
- Runs k-means and interprets the clusters
- Chooses the number of clusters with the elbow method or silhouette
- Knows when k-means is unsuitable
Sign in to do the exercises and build your mastery up.
Sources
- scikit-learn User Guide (BSD-3) — BSD-3-Clause
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0