Skip to content
AI-grafen
DAI developerMathematics· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Exponential functions

Be able to describe growth and decay with exponential functions and connect them to e^x.

Practise in Mattegrafen ↗ · Matematik 1c

Prerequisites

Intuition

Linear: you add the same amount each time. 100, 110, 120, 130 …

Exponential: you multiply by the same amount each time. 100, 110, 121, 133.1 …

The difference looks small at first and then becomes gigantic.

StepLinear (+10)Exponential (×1.1)
0100100
10200259
5060011 739
1001 1001 378 061

The form is f(x)=a⋅bxf(x) = a \cdot b^x where bb is the growth factor:

bbWhat happens
b>1b > 1it grows
b=1b = 1it stands still
0<b<10 < b < 1it decays towards zero, but never reaches it

The number e≈2.718e \approx 2.718 is the choice of base that makes the derivative as simple as possible: ddxex=ex\frac{d}{dx}e^x = e^x. The function is its own derivative. That is why ee turns up everywhere in the mathematics behind ML.

Formal

Half-life and doubling time. If something decays by the factor bb per step:

half-life=ln⁡0.5ln⁡b\text{half-life} = \frac{\ln 0.5}{\ln b}

For growth the same formula is used with ln⁡2\ln 2. The rule of thumb «the rule of 72» is an approximation of exactly this: at pp per cent growth per period the doubling time is roughly 72/p72/p periods.

Where exponential functions turn up in ML:

SettingExpressionWhy
A learning-rate scheduleηt=η0⋅γt\eta_t = \eta_0 \cdot \gamma^tlarge steps first, small ones later
A moving average (Adam)mt=βmt−1+(1−β)gtm_t = \beta m_{t-1} + (1-\beta)g_told gradients lose weight exponentially
Softmaxezi/∑jezje^{z_i} / \sum_j e^{z_j}turns numbers into probabilities, amplifies the differences
Discounting in RLγt\gamma^tfuture rewards weigh less
The forgetting curvee−t/τe^{-t/\tau}knowledge fades without repetition

The last row is AI-grafen's own: the platform's mastery model lets mastery decay exponentially with time, with a time constant τ\tau that grows every time you revise. That is why the second revision lasts longer than the first.

A numerical trap. e1000e^{1000} overflows in floating point. That is why you compute with logarithms instead — and why softmax always normalises by subtracting the largest value first:

ezi∑jezj=ezi−max⁡z∑jezj−max⁡z\frac{e^{z_i}}{\sum_j e^{z_j}} = \frac{e^{z_i - \max z}}{\sum_j e^{z_j - \max z}}

The same answer, but no overflows. It is one of the most common numerical tricks in the whole field.

Code

import math

# Linear against exponential
for step in (0, 10, 50, 100):
    print(f"{step:4d}  linear {100 + 10 * step:>10,}   exp {100 * 1.1 ** step:>15,.0f}")

# e is its own derivative — check it numerically
h = 1e-7
for x in (0.0, 1.0, 2.0):
    derivative = (math.exp(x + h) - math.exp(x)) / h
    print(f"x={x}  f(x)={math.exp(x):.5f}  f'(x)≈{derivative:.5f}")

# The half-life at a 5 % decrease per step
b = 0.95
print(round(math.log(0.5) / math.log(b), 1), "steps")     # 13.5 steps

# A learning-rate schedule
eta0, gamma = 0.1, 0.95
for epoch in (0, 10, 50, 100):
    print(f"epoch {epoch:3d}: lr = {eta0 * gamma ** epoch:.6f}")

# The numerical trap — and the solution
z = [1000.0, 1001.0, 1002.0]
try:
    naive = [math.exp(v) for v in z]
except OverflowError as e:
    print("naive softmax:", e)                # math range error

m = max(z)
exp_stable = [math.exp(v - m) for v in z]
s = sum(exp_stable)
print([round(v / s, 4) for v in exp_stable])   # [0.0900, 0.2447, 0.6652]

The last lines show the trick in its purest form: subtracting the maximum does not change the answer (the factor cancels) but removes the overflow entirely.

Mastery means

  • Tells exponential from linear change
  • Uses e^x and interprets the growth factor
  • Recognises exponential decay in ML settings

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences