Exponential functions
Be able to describe growth and decay with exponential functions and connect them to e^x.
Practise in Mattegrafen ↗ · Matematik 1cPrerequisites
Intuition
Linear: you add the same amount each time. 100, 110, 120, 130 …
Exponential: you multiply by the same amount each time. 100, 110, 121, 133.1 …
The difference looks small at first and then becomes gigantic.
| Step | Linear (+10) | Exponential (×1.1) |
|---|---|---|
| 0 | 100 | 100 |
| 10 | 200 | 259 |
| 50 | 600 | 11 739 |
| 100 | 1 100 | 1 378 061 |
The form is where is the growth factor:
| What happens | |
|---|---|
| it grows | |
| it stands still | |
| it decays towards zero, but never reaches it |
The number is the choice of base that makes the derivative as simple as possible: . The function is its own derivative. That is why turns up everywhere in the mathematics behind ML.
Formal
Half-life and doubling time. If something decays by the factor per step:
For growth the same formula is used with . The rule of thumb «the rule of 72» is an approximation of exactly this: at per cent growth per period the doubling time is roughly periods.
Where exponential functions turn up in ML:
| Setting | Expression | Why |
|---|---|---|
| A learning-rate schedule | large steps first, small ones later | |
| A moving average (Adam) | old gradients lose weight exponentially | |
| Softmax | turns numbers into probabilities, amplifies the differences | |
| Discounting in RL | future rewards weigh less | |
| The forgetting curve | knowledge fades without repetition |
The last row is AI-grafen's own: the platform's mastery model lets mastery decay exponentially with time, with a time constant that grows every time you revise. That is why the second revision lasts longer than the first.
A numerical trap. overflows in floating point. That is why you compute with logarithms instead — and why softmax always normalises by subtracting the largest value first:
The same answer, but no overflows. It is one of the most common numerical tricks in the whole field.
Code
import math
# Linear against exponential
for step in (0, 10, 50, 100):
print(f"{step:4d} linear {100 + 10 * step:>10,} exp {100 * 1.1 ** step:>15,.0f}")
# e is its own derivative — check it numerically
h = 1e-7
for x in (0.0, 1.0, 2.0):
derivative = (math.exp(x + h) - math.exp(x)) / h
print(f"x={x} f(x)={math.exp(x):.5f} f'(x)≈{derivative:.5f}")
# The half-life at a 5 % decrease per step
b = 0.95
print(round(math.log(0.5) / math.log(b), 1), "steps") # 13.5 steps
# A learning-rate schedule
eta0, gamma = 0.1, 0.95
for epoch in (0, 10, 50, 100):
print(f"epoch {epoch:3d}: lr = {eta0 * gamma ** epoch:.6f}")
# The numerical trap — and the solution
z = [1000.0, 1001.0, 1002.0]
try:
naive = [math.exp(v) for v in z]
except OverflowError as e:
print("naive softmax:", e) # math range error
m = max(z)
exp_stable = [math.exp(v - m) for v in z]
s = sum(exp_stable)
print([round(v / s, 4) for v in exp_stable]) # [0.0900, 0.2447, 0.6652]
The last lines show the trick in its purest form: subtracting the maximum does not change the answer (the factor cancels) but removes the overflow entirely.
Mastery means
- Tells exponential from linear change
- Uses e^x and interprets the growth factor
- Recognises exponential decay in ML settings
Sign in to do the exercises and build your mastery up.
Sources
- Matteboken (Mattecentrum) — free to read, non-profit association
- Khan Academy — matematik — CC BY-NC-SA 3.0
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0