Logarithms
Be able to use logarithms to solve exponential equations and understand log scales — the basis for log loss and log probabilities.
Practise in Mattegrafen ↗ · Matematik 2cPrerequisites
Intuition
The logarithm answers the question: «raised to what?»
It is therefore the opposite of the exponential function — just as subtraction is the opposite of addition.
| Base | Written | Used for |
|---|---|---|
| 10 | or | decibels, pH, orders of magnitude |
| all mathematical analysis, ML | ||
| 2 | information, bits, entropy |
The logarithm's superpower: it turns multiplication into addition.
That sounds like a curiosity but is decisive in practice: multiply 500 probabilities (each smaller than 1) and you get a number so small that the computer rounds it to zero. Add their logarithms and you get a manageable negative number.
That is why all machine learning computes in log space.
Formal
The laws — all of them follow from the laws of exponents:
| Law | |
|---|---|
| a product → a sum | |
| a quotient → a difference | |
| a power → a factor | |
| , | |
| change of base |
Solving exponential equations. Take the logarithm of both sides:
Why ML computes in log space — three reasons:
- Numerical stability. is zero in floating point. is not.
- Products become sums. The probability of a whole sequence is a product over the tokens; the log probability is a sum, which can moreover be differentiated term by term.
- Log loss is the natural loss. Maximising the log probability of the right answer is the same thing as minimising the cross-entropy:
For a confident and correct guess (), . For a confident and wrong guess () it goes to infinity. Log loss therefore punishes confident stupidity infinitely hard — which is precisely what you want.
Perplexity is the same thing in a more readable wrapping: , interpreted as «how many alternatives the model is effectively choosing between». If the log loss goes from 2.3 to 2.0 it sounds small — but the perplexity goes from 10.0 to 7.4, which is a quarter off the model's effective uncertainty.
Code
import math
print(math.log2(8), math.log10(1000), round(math.log(math.e), 4)) # 3.0 3.0 1.0
# Solving 3 · 2^x = 96
print(math.log(96 / 3) / math.log(2)) # 5.0
# Why log space is needed: 400 probabilities multiplied
p = [0.1] * 400
print(math.prod(p)) # 0.0 ← all the information gone
print(sum(math.log(x) for x in p)) # -921.03 ← still exact
# Log loss punishes confident errors
for p_correct in (0.99, 0.9, 0.5, 0.1, 0.01, 1e-8):
print(f" p={p_correct:<8} log loss={-math.log(p_correct):.3f}")
# p=0.99 log loss=0.010
# p=0.5 log loss=0.693
# p=0.01 log loss=4.605
# p=1e-08 log loss=18.421 ← a confident error costs enormously
# Perplexity makes log loss readable
for loss in (2.3, 2.0, 1.6):
print(f" loss={loss} perplexity={math.exp(loss):.1f}")
# loss=2.3 perplexity=10.0
# loss=2.0 perplexity=7.4
# loss=1.6 perplexity=5.0
# log1p is more exact than log(1+x) for small x
x = 1e-15
print(math.log(1 + x), math.log1p(x)) # 1.110223e-15 1.0e-15 ← log1p is exact
Mastery means
- Uses the laws of logarithms
- Solves exponential equations with a logarithm
- Explains why ML computes in log space
Sign in to do the exercises and build your mastery up.
Sources
- Matteboken (Mattecentrum) — free to read, non-profit association
- Khan Academy — matematik — CC BY-NC-SA 3.0
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0