Nearest neighbour: classification based on similarity
Be able to classify a new observation based on which existing observations it most resembles.
Prerequisites
Everyday explanation
You have a dataset with points. Blue points represent 'customers who bought', red are 'customers who declined'. Now a new customer arrives. Which category do they belong to?
Nearest neighbour: Look at the existing customer closest to the new one in the data space. Did that blue customer buy? Then we guess the new one will buy too.
It is that simple. No complex model, no difficult formulas — just the principle that similar objects behave similarly. It is one of the simplest and most intuitive methods in machine learning and often works well as a baseline.
Interactive
Try it yourself. Imagine a grid from 0–10 on both axes. x = age, y = income.
- Buyers (blue): (8, 9), (7, 7), (9, 8)
- Decliners (red): (3, 2), (2, 4), (4, 3)
- New customer: (6, 6). Which group is closest?
Measure the distance. (7, 7) is closest → Buyer.
More neighbours give better stability. Look at the three nearest instead of one. If two of three are blue: Buyer. This protects against single outliers that happened to land in the wrong place.
Add an 'outlier' who declined at (8, 8) and test again with one neighbour versus three. Do you see the difference?
Mastery means
- Classifies a new observation based on nearest neighbours
- Explains why the number of neighbours affects robustness
Sign in to do the exercises and build your mastery up.
Sources
- CS Unplugged (CC BY-SA 4.0) — CC BY-SA 4.0
- Wikipedia — k-nearest neighbors algorithm (CC BY-SA 4.0) — CC BY-SA 4.0