Skip to content
AI-grafen
EUniversityMathematics· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

PCA — principal component analysis

Be able to reduce dimensions with PCA and interpret the explained variance.

Prerequisites

Intuition

PCA finds the directions in which the data varies the most, and projects onto them.

Imagine a cloud of points stretched out along a diagonal. The first principal component is that line. The second is perpendicular to it, along the second largest spread. And so on.

Why variation? Because a direction in which everything lies in the same place does not tell the data points apart — it carries no information about what is different.

Three things PCA is used for:

UseHow
Compression100 columns → 10, with 95 % of the variation left
Visualisationproject down to 2D and plot
Denoisingthrow away the last components, which are often measurement noise

The decisive preparation: always centre the data (subtract the mean), and scale if the columns have different units. Otherwise PCA measures units instead of structure.

Formal

Two equivalent descriptions:

  1. The eigenvectors of the covariance matrix C=1n−1X⊤XC = \frac{1}{n-1}X^\top X (after centring), sorted by eigenvalue.
  2. The right singular vectors VV from the SVD of the centred XX.

The second is what is actually computed — it is numerically more stable and never needs to form the covariance matrix. The relationship is λi=σi2/(n−1)\lambda_i = \sigma_i^2/(n-1).

The explained variance for component ii:

λi∑jλj=σi2∑jσj2\frac{\lambda_i}{\sum_j \lambda_j} = \frac{\sigma_i^2}{\sum_j \sigma_j^2}

How many components? Three common rules:

RuleHow
A thresholdkeep going until the cumulative variance reaches 90–95 %
The elbowlook for the knee in the scree plot
Downstreamchoose the one that gives the best result in the model that follows

The last is the most honest when PCA is a preprocessing step.

Scaling is not optional. If you have weight in grams (0–100 000) and length in metres (0–2), the first principal component will be weight, whatever the data is really about. Variance is not dimensionless. Standardise first, unless all the columns already have the same unit.

Leakage: fit on the training data, transform on both training and test. Running fit_transform on everything is the same error as with scaling.

Limitations:

LimitationMeans
Lineardoes not capture curved manifolds — a spiral is not unrolled
Variance ≠ usefulnessthe most important direction for classification can have low variance
Hard-to-interpret componentsevery component is a mixture of all the original variables
Sensitive to outliersone extreme point can dominate the first component

The second point is important: PCA is unsupervised and knows nothing about the labels. If you need the direction that best separates the classes, LDA is the right tool, not PCA.

For visualisation t-SNE and UMAP are often better at showing cluster structure — but they do not preserve distances globally and their output should not be fed onwards into a model.

Code

import numpy as np

# 800 points in 30 dimensions — but really only 4 latent factors plus noise
rng = np.random.default_rng(0)
Z = rng.normal(size=(800, 4))
X = Z @ rng.normal(size=(4, 30)) + 0.5 * rng.normal(size=(800, 30))

Xc = X - X.mean(axis=0)                      # CENTRE — otherwise PCA measures the wrong thing
U, s, Vt = np.linalg.svd(Xc, full_matrices=False)
share = s**2 / (s**2).sum()

print(np.round(s[:8], 1))
# [196.9 152.2 134.3 120.5  16.6  16.1  15.9  15.7]
#   ↑ four large ones, then a clear jump down to the noise level

print(np.round(share[:6], 4))                # [0.3892 0.2324 0.1811 0.1458 0.0028 0.0026]
print(np.round(np.cumsum(share)[:6], 4))     # [0.3892 0.6215 0.8026 0.9484 0.9511 0.9537]
print("components for 95 %:", int(np.searchsorted(np.cumsum(share), 0.95) + 1))   # 5

# The reconstruction error: how much is lost with k components?
for k in (2, 4, 10):
    Xk = U[:, :k] * s[:k] @ Vt[:k]
    print(f"  k={k:>2}  relative error {np.linalg.norm(Xc - Xk) / np.linalg.norm(Xc):.4f}")
#   k= 2  relative error 0.6152
#   k= 4  relative error 0.2273      ← the four real factors
#   k=10  relative error 0.1913      ← barely better; the rest is noise

# What the scaling does: one column in a different unit
Xo = X.copy(); Xo[:, 0] *= 10_000

def shares(M, n=2):
    s2 = np.linalg.svd(M - M.mean(axis=0), compute_uv=False) ** 2
    return np.round((s2 / s2.sum())[:n], 3)

print("without scaling:", shares(Xo))                                  # [1. 0.]
print("with scaling:   ", shares((Xo - Xo.mean(0)) / Xo.std(0)))       # [0.332 0.247]
#  ↑ without standardisation the large column eats the whole first component

In practice you put PCA in a Pipeline together with the scaling and the model, so that fit is rerun within each cross-validation fold:

from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score

pipe = Pipeline([("scale", StandardScaler()),
                 ("pca", PCA(n_components=20)),
                 ("classifier", LogisticRegression(max_iter=2000))])
print(cross_val_score(pipe, X, y, cv=5).mean())

Vary n_components and let the result decide — that is the most honest rule when PCA is a preprocessing step. A selection of components often beats all of them, since the last ones carry the most noise.

Mastery means

  • Performs PCA and interprets the explained variance
  • Knows why centring and scaling are needed
  • Knows PCA's limitations

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences