Skip to content
AI-grafen
EUniversityMathematics· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Eigenvalues and eigenvectors

Be able to compute eigenvalues for small matrices and explain what they say about a map.

Prerequisites

Intuition

A matrix A is a map. For most vectors both the direction and the length change. But some directions keep their direction and are only scaled:

Av=λvAv = \lambda v

Then v is an eigenvector and λ its eigenvalue.

For [[2,0],[0,3]] the x-axis and the y-axis are eigendirections with the eigenvalues 2 and 3: the map stretches x twofold and y threefold. For a rotation matrix there are no real eigenvectors — everything is turned.

Why it matters in ML: the product of many matrices (deep networks, RNNs over time) is dominated by the largest eigenvalue. If it is > 1 the signal grows exponentially with the depth; if it is < 1 it dies out.

Formal

Computing them for 2×2: solve the characteristic equation det⁡(A−λI)=0\det(A - \lambda I) = 0.

For A=(4123)A = \begin{pmatrix}4 & 1\\ 2 & 3\end{pmatrix}: (4−λ)(3−λ)−2=λ2−7λ+10=0⇒λ=5, 2(4-\lambda)(3-\lambda) - 2 = \lambda^2 - 7\lambda + 10 = 0 \Rightarrow \lambda = 5,\ 2

The eigenvector for λ=5\lambda = 5: solve (A−5I)v=0⇒v=(1,1)(A - 5I)v = 0 \Rightarrow v = (1, 1).

Useful relationships:

  • ∑iλi=tr(A)\sum_i \lambda_i = \text{tr}(A) (the trace) — a quick check: 5 + 2 = 4 + 3 ✔
  • ∏iλi=det⁡(A)\prod_i \lambda_i = \det(A) — 5 · 2 = 12 − 2 ✔
  • The spectral radius ρ(A)=max⁡∣λi∣\rho(A) = \max|\lambda_i| governs whether AnA^n grows or shrinks.

Symmetric matrices (such as covariance matrices and Hessians) always have real eigenvalues and orthogonal eigenvectors — the spectral theorem. That is why PCA works: the eigenvectors of the covariance matrix are orthogonal directions sorted by variance.

The Hessian's eigenvalues at a critical point decide whether it is a minimum (all positive), a maximum (all negative) or a saddle point (mixed signs) — and the ratio between the largest and the smallest eigenvalue (the condition number) decides how badly gradient descent zigzags.

Code

import numpy as np

A = np.array([[4.0, 1.0], [2.0, 3.0]])
values, vectors = np.linalg.eig(A)
print(values.round(3))                       # [5. 2.]
print(vectors.round(3))                      # the columns are the eigenvectors
print(np.trace(A), values.sum())             # 7.0 7.0   ← a check
print(round(np.linalg.det(A), 3), round(values.prod(), 3))   # 10.0 10.0

# Verify the definition for the first eigenpair
v = vectors[:, 0]
print(np.allclose(A @ v, values[0] * v))     # True

# The spectral radius governs whether repeated application grows or dies
for scale in (0.9, 1.0, 1.1):
    M = A * scale / np.abs(values).max()
    print(scale, round(float(np.abs(np.linalg.eigvals(np.linalg.matrix_power(M, 50))).max()), 6))
# 0.9  0.005154   ← dies out
# 1.0  1.0
# 1.1  117.39     ← explodes

Mastery means

  • Computes the eigenvalues of a 2×2 matrix
  • Interprets eigenvalues as scale factors in the eigendirections
  • Connects the spectrum to stability

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences