Quadratic functions
Be able to draw parabolas, find the vertex and solve quadratic equations — the shape of a simple loss surface.
Practise in Mattegrafen ↗ · Matematik 2cPractise in Mattegrafen ↗ · Metoder för att lösa andragradsekvationePrerequisites
Intuition
draws a parabola — a bowl.
| If | Then |
|---|---|
| the bowl opens upwards, has a minimum | |
| the bowl opens downwards, has a maximum | |
| large $ | a |
| small $ | a |
The vertex (the bottom or the top of the bowl) lies at
That formula is worth memorising. It follows directly from the symmetry: the parabola is mirror-symmetric about a vertical line through the vertex, so the vertex lies midway between the roots.
Why AI cares: the mean squared error of a linear model is a parabola in each individual parameter. When you hear that «the model is looking for the minimum of the loss function» — this bowl is the one whose bottom it is looking for.
Formal
The roots — where :
The discriminant decides how many there are:
| Number of roots | The graph | |
|---|---|---|
| two | crosses the x-axis in two places | |
| one (a double root) | touches the x-axis | |
| none (real) | lies entirely above or entirely below |
Completing the square rewrites the function so that the vertex is visible directly:
That is the form that makes it obvious that the minimum is and is reached at : the square is always ≥ 0, and is zero precisely there.
The connection to machine learning, concretely. Fit to data with the mean squared error:
That is a parabola in , with . The minimum lies at — and that is exactly the solution least squares gives.
That is why a «convex loss» is something you want: a bowl has one bottom, and gradient descent cannot get stuck anywhere else.
Code
import numpy as np
def vertex(a, b, c):
x = -b / (2 * a)
return x, a * x ** 2 + b * x + c
def roots(a, b, c):
D = b ** 2 - 4 * a * c
if D < 0:
return []
if D == 0:
return [-b / (2 * a)]
r = D ** 0.5
return [(-b - r) / (2 * a), (-b + r) / (2 * a)]
print(vertex(1, -4, 3)) # (2.0, -1.0)
print(roots(1, -4, 3)) # [1.0, 3.0] — the vertex lies midway between them
print(roots(1, 2, 5)) # [] — D = 4 - 20 < 0
# The loss function for y = w·x IS a parabola in w
x = np.array([1.0, 2.0, 3.0, 4.0])
y = np.array([2.1, 3.9, 6.2, 7.8])
def L(w):
return float(np.mean((y - w * x) ** 2))
a = float(np.mean(x ** 2))
b = float(-2 * np.mean(x * y))
c = float(np.mean(y ** 2))
w_star = -b / (2 * a)
print(round(w_star, 4), round(L(w_star), 5)) # 1.99 0.02425
# The same answer least squares gives directly
print(round(float(np.sum(x * y) / np.sum(x ** 2)), 4)) # 1.99
# And the parabola really is a bowl: every other w is worse
for w in (1.5, 1.9, w_star, 2.1, 2.5):
print(f" w={w:.4f} L={L(w):.5f}")
# w=1.5000 L=1.82500
# w=1.9000 L=0.08500
# w=1.9900 L=0.02425 ← the bottom
# w=2.1000 L=0.11500
# w=2.5000 L=1.97500
The last lines are the whole point: training a linear model is finding the bottom of a parabola, and for this particular model the answer can be computed directly instead of searched for.
Mastery means
- Draws a parabola and finds its vertex
- Solves quadratic equations
- Connects the shape of the parabola to a loss surface
Sign in to do the exercises and build your mastery up.
Sources
- Matteboken (Mattecentrum) — free to read, non-profit association
- Khan Academy — matematik — CC BY-NC-SA 3.0
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0