Dimensionality reduction: PCA, t-SNE, UMAP
Be able to reduce dimensions for visualisation and interpret the result critically.
Prerequisites
- DClustering: k-means and hierarchicalrequired
- EPCA — principal component analysisrequired
Intuition
Three methods, three different purposes — they are not interchangeable.
| Method | Preserves | Used for | Deterministic |
|---|---|---|---|
| PCA | global distances, linearly | preprocessing, compression | yes |
| t-SNE | local neighbourhoods | visualisation | no |
| UMAP | local plus some global | visualisation, sometimes features | no (but more stable) |
The decisive difference: PCA's output can be fed onwards into a model. t-SNE's cannot — it has no transform for new points, the result changes between runs, and the distances in the picture do not mean what they look as though they mean.
t-SNE and UMAP are for the eye. They make beautiful clusters, and it is easy to read more into the picture than is there.
Formal
What t-SNE actually does: it defines probabilities that points are neighbours in the high dimension, does the same in the low dimension (with a heavier tail, Student's t), and minimises the KL divergence between the two distributions.
The heavy tail is what creates the clear clusters — it allows points that are not neighbours to be pushed far apart.
Four things t-SNE pictures do NOT say — and which are nearly always over-interpreted:
| What people think | What holds |
|---|---|
| The cluster size means something | no — t-SNE expands sparse clusters and compresses dense ones |
| The distances between clusters mean something | no — global distances are not preserved |
| The picture is stable | no — different seeds give different pictures |
| Clusters in the picture are real clusters | not necessarily — t-SNE can create clusters in pure random data |
The last is worth trying yourself: run t-SNE on Gaussian noise and you often get something that looks like structure.
Perplexity is t-SNE's most important parameter (typically 5–50) and governs how many neighbours count as local. A low perplexity gives many small clusters, a high one fewer large ones. Always run several values — if the structure looks the same at 5, 30 and 50 it is probably real.
UMAP builds on a graph-theoretic formulation, is faster, scales better and preserves more global structure. It also has a transform for new points. But the same warnings mostly apply: cluster distances should not be over-interpreted.
A practical working order:
- PCA first to 30–50 dimensions. That removes noise and makes t-SNE/UMAP dramatically faster.
- Run t-SNE or UMAP on the PCA output.
- Colour by a known variable (a label, time, a source). Without colour the picture is nearly always uninterpretable.
- Verify with several seeds and parameters before drawing a conclusion.
- Never use t-SNE coordinates as features in a model.
Point 3 is what makes the visualisation useful: a picture where the clusters coincide with a variable you did not feed in is a genuine finding.
Code
import numpy as np
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
rng = np.random.default_rng(0)
# 1. The working order: PCA first, then t-SNE
def visualise(X, n_pca=50, perplexity=30, seed=0):
Xp = PCA(n_components=min(n_pca, X.shape[1]), random_state=seed).fit_transform(X)
return TSNE(n_components=2, perplexity=perplexity, random_state=seed,
init="pca").fit_transform(Xp)
# 2. A WARNING: t-SNE finds clusters even in pure noise. Try it yourself.
noise = rng.normal(size=(500, 50))
embedding = visualise(noise, perplexity=30)
print("t-SNE on pure noise gave a picture with visible structure — it means nothing")
# 3. A stability check: the same data, different seeds
def stability(X, seeds=(0, 1, 2), perplexity=30):
"""How similar do the neighbourhoods become between runs? 1.0 = identical."""
from sklearn.neighbors import NearestNeighbors
neighbours = []
for s in seeds:
E = visualise(X, perplexity=perplexity, seed=s)
nn = NearestNeighbors(n_neighbors=11).fit(E)
neighbours.append([set(r[1:]) for r in nn.kneighbors(E, return_distance=False)])
overlap = [len(a & b) / 10 for g1, g2 in zip(neighbours, neighbours[1:])
for a, b in zip(g1, g2)]
return round(float(np.mean(overlap)), 3)
# 4. A parameter check: the same structure at different perplexities?
for p in (5, 30, 50):
E = visualise(X, perplexity=p)
print(f" perplexity={p}: run it and compare the pictures visually")
# 5. PCA is the only one of the three that may be fed onwards into a model
p = PCA(n_components=20).fit(X_train) # fit on the TRAINING data
Z_train = p.transform(X_train)
Z_test = p.transform(X_test) # the same transform on the test — t-SNE cannot do this
The block at point 2 is the most important thing in this whole node. Run t-SNE on 500 random points in 50 dimensions and look at the result. That the picture looks structured even though no structure exists is the best vaccination against over-interpretation there is.
Mastery means
- Chooses the method according to the purpose
- Interprets t-SNE and UMAP critically
- Knows which properties are not preserved
Sign in to do the exercises and build your mastery up.
Sources
- Wattenberg m.fl. — How to Use t-SNE Effectively (Distill, CC BY 4.0) — CC BY 4.0
- arXiv — UMAP: Uniform Manifold Approximation and Projection — arXiv (open access; licence per article)
- scikit-learn User Guide (BSD-3) — BSD-3-Clause