Skip to content
AI-grafen
EUniversityClassical machine learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Dimensionality reduction: PCA, t-SNE, UMAP

Be able to reduce dimensions for visualisation and interpret the result critically.

Prerequisites

Intuition

Three methods, three different purposes — they are not interchangeable.

MethodPreservesUsed forDeterministic
PCAglobal distances, linearlypreprocessing, compressionyes
t-SNElocal neighbourhoodsvisualisationno
UMAPlocal plus some globalvisualisation, sometimes featuresno (but more stable)

The decisive difference: PCA's output can be fed onwards into a model. t-SNE's cannot — it has no transform for new points, the result changes between runs, and the distances in the picture do not mean what they look as though they mean.

t-SNE and UMAP are for the eye. They make beautiful clusters, and it is easy to read more into the picture than is there.

Formal

What t-SNE actually does: it defines probabilities that points are neighbours in the high dimension, does the same in the low dimension (with a heavier tail, Student's t), and minimises the KL divergence between the two distributions.

The heavy tail is what creates the clear clusters — it allows points that are not neighbours to be pushed far apart.

Four things t-SNE pictures do NOT say — and which are nearly always over-interpreted:

What people thinkWhat holds
The cluster size means somethingno — t-SNE expands sparse clusters and compresses dense ones
The distances between clusters mean somethingno — global distances are not preserved
The picture is stableno — different seeds give different pictures
Clusters in the picture are real clustersnot necessarily — t-SNE can create clusters in pure random data

The last is worth trying yourself: run t-SNE on Gaussian noise and you often get something that looks like structure.

Perplexity is t-SNE's most important parameter (typically 5–50) and governs how many neighbours count as local. A low perplexity gives many small clusters, a high one fewer large ones. Always run several values — if the structure looks the same at 5, 30 and 50 it is probably real.

UMAP builds on a graph-theoretic formulation, is faster, scales better and preserves more global structure. It also has a transform for new points. But the same warnings mostly apply: cluster distances should not be over-interpreted.

A practical working order:

  1. PCA first to 30–50 dimensions. That removes noise and makes t-SNE/UMAP dramatically faster.
  2. Run t-SNE or UMAP on the PCA output.
  3. Colour by a known variable (a label, time, a source). Without colour the picture is nearly always uninterpretable.
  4. Verify with several seeds and parameters before drawing a conclusion.
  5. Never use t-SNE coordinates as features in a model.

Point 3 is what makes the visualisation useful: a picture where the clusters coincide with a variable you did not feed in is a genuine finding.

Code

import numpy as np
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE

rng = np.random.default_rng(0)

# 1. The working order: PCA first, then t-SNE
def visualise(X, n_pca=50, perplexity=30, seed=0):
    Xp = PCA(n_components=min(n_pca, X.shape[1]), random_state=seed).fit_transform(X)
    return TSNE(n_components=2, perplexity=perplexity, random_state=seed,
                init="pca").fit_transform(Xp)

# 2. A WARNING: t-SNE finds clusters even in pure noise. Try it yourself.
noise = rng.normal(size=(500, 50))
embedding = visualise(noise, perplexity=30)
print("t-SNE on pure noise gave a picture with visible structure — it means nothing")

# 3. A stability check: the same data, different seeds
def stability(X, seeds=(0, 1, 2), perplexity=30):
    """How similar do the neighbourhoods become between runs? 1.0 = identical."""
    from sklearn.neighbors import NearestNeighbors
    neighbours = []
    for s in seeds:
        E = visualise(X, perplexity=perplexity, seed=s)
        nn = NearestNeighbors(n_neighbors=11).fit(E)
        neighbours.append([set(r[1:]) for r in nn.kneighbors(E, return_distance=False)])
    overlap = [len(a & b) / 10 for g1, g2 in zip(neighbours, neighbours[1:])
               for a, b in zip(g1, g2)]
    return round(float(np.mean(overlap)), 3)

# 4. A parameter check: the same structure at different perplexities?
for p in (5, 30, 50):
    E = visualise(X, perplexity=p)
    print(f"  perplexity={p}: run it and compare the pictures visually")

# 5. PCA is the only one of the three that may be fed onwards into a model
p = PCA(n_components=20).fit(X_train)     # fit on the TRAINING data
Z_train = p.transform(X_train)
Z_test = p.transform(X_test)              # the same transform on the test — t-SNE cannot do this

The block at point 2 is the most important thing in this whole node. Run t-SNE on 500 random points in 50 dimensions and look at the result. That the picture looks structured even though no structure exists is the best vaccination against over-interpretation there is.

Mastery means

  • Chooses the method according to the purpose
  • Interprets t-SNE and UMAP critically
  • Knows which properties are not preserved

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences