PCA — principal component analysis
Be able to reduce dimensions with PCA and interpret the explained variance.
Prerequisites
Intuition
PCA finds the directions in which the data varies the most, and projects onto them.
Imagine a cloud of points stretched out along a diagonal. The first principal component is that line. The second is perpendicular to it, along the second largest spread. And so on.
Why variation? Because a direction in which everything lies in the same place does not tell the data points apart — it carries no information about what is different.
Three things PCA is used for:
| Use | How |
|---|---|
| Compression | 100 columns → 10, with 95 % of the variation left |
| Visualisation | project down to 2D and plot |
| Denoising | throw away the last components, which are often measurement noise |
The decisive preparation: always centre the data (subtract the mean), and scale if the columns have different units. Otherwise PCA measures units instead of structure.
Formal
Two equivalent descriptions:
- The eigenvectors of the covariance matrix (after centring), sorted by eigenvalue.
- The right singular vectors from the SVD of the centred .
The second is what is actually computed — it is numerically more stable and never needs to form the covariance matrix. The relationship is .
The explained variance for component :
How many components? Three common rules:
| Rule | How |
|---|---|
| A threshold | keep going until the cumulative variance reaches 90–95 % |
| The elbow | look for the knee in the scree plot |
| Downstream | choose the one that gives the best result in the model that follows |
The last is the most honest when PCA is a preprocessing step.
Scaling is not optional. If you have weight in grams (0–100 000) and length in metres (0–2), the first principal component will be weight, whatever the data is really about. Variance is not dimensionless. Standardise first, unless all the columns already have the same unit.
Leakage: fit on the training data, transform on both training and test. Running fit_transform on everything is the same error as with scaling.
Limitations:
| Limitation | Means |
|---|---|
| Linear | does not capture curved manifolds — a spiral is not unrolled |
| Variance ≠ usefulness | the most important direction for classification can have low variance |
| Hard-to-interpret components | every component is a mixture of all the original variables |
| Sensitive to outliers | one extreme point can dominate the first component |
The second point is important: PCA is unsupervised and knows nothing about the labels. If you need the direction that best separates the classes, LDA is the right tool, not PCA.
For visualisation t-SNE and UMAP are often better at showing cluster structure — but they do not preserve distances globally and their output should not be fed onwards into a model.
Code
import numpy as np
# 800 points in 30 dimensions — but really only 4 latent factors plus noise
rng = np.random.default_rng(0)
Z = rng.normal(size=(800, 4))
X = Z @ rng.normal(size=(4, 30)) + 0.5 * rng.normal(size=(800, 30))
Xc = X - X.mean(axis=0) # CENTRE — otherwise PCA measures the wrong thing
U, s, Vt = np.linalg.svd(Xc, full_matrices=False)
share = s**2 / (s**2).sum()
print(np.round(s[:8], 1))
# [196.9 152.2 134.3 120.5 16.6 16.1 15.9 15.7]
# ↑ four large ones, then a clear jump down to the noise level
print(np.round(share[:6], 4)) # [0.3892 0.2324 0.1811 0.1458 0.0028 0.0026]
print(np.round(np.cumsum(share)[:6], 4)) # [0.3892 0.6215 0.8026 0.9484 0.9511 0.9537]
print("components for 95 %:", int(np.searchsorted(np.cumsum(share), 0.95) + 1)) # 5
# The reconstruction error: how much is lost with k components?
for k in (2, 4, 10):
Xk = U[:, :k] * s[:k] @ Vt[:k]
print(f" k={k:>2} relative error {np.linalg.norm(Xc - Xk) / np.linalg.norm(Xc):.4f}")
# k= 2 relative error 0.6152
# k= 4 relative error 0.2273 ← the four real factors
# k=10 relative error 0.1913 ← barely better; the rest is noise
# What the scaling does: one column in a different unit
Xo = X.copy(); Xo[:, 0] *= 10_000
def shares(M, n=2):
s2 = np.linalg.svd(M - M.mean(axis=0), compute_uv=False) ** 2
return np.round((s2 / s2.sum())[:n], 3)
print("without scaling:", shares(Xo)) # [1. 0.]
print("with scaling: ", shares((Xo - Xo.mean(0)) / Xo.std(0))) # [0.332 0.247]
# ↑ without standardisation the large column eats the whole first component
In practice you put PCA in a Pipeline together with the scaling and the model, so that fit is rerun within each cross-validation fold:
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
pipe = Pipeline([("scale", StandardScaler()),
("pca", PCA(n_components=20)),
("classifier", LogisticRegression(max_iter=2000))])
print(cross_val_score(pipe, X, y, cv=5).mean())
Vary n_components and let the result decide — that is the most honest rule when PCA is a preprocessing step. A selection of components often beats all of them, since the last ones carry the most noise.
Mastery means
- Performs PCA and interprets the explained variance
- Knows why centring and scaling are needed
- Knows PCA's limitations
Sign in to do the exercises and build your mastery up.
Sources
- Mathematics for Machine Learning (Deisenroth m.fl.) — free to read online (authors' edition)
- scikit-learn User Guide (BSD-3) — BSD-3-Clause
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0