Bayes' theorem
Be able to update a probability with new evidence and apply Bayes to classification and diagnostics.
Prerequisites
- DConditional probabilityrequired
Intuition
Bayes' theorem is a rule for changing your mind when you learn something new.
| Part | Name | Means |
|---|---|---|
| the prior | what I believed beforehand | |
| the likelihood | how expected the evidence is if A is true | |
| the posterior | what I believe afterwards | |
| the evidence | how expected the evidence is at all |
The decisive insight: and are not the same thing.
- «If it is raining the ground is wet» — nearly always true.
- «If the ground is wet it is raining» — often false. Somebody has been hosing it.
Mixing them up is probably the most common probability error there is, and it has a name: confusion of the inverse.
Formal
The classic example, in numbers. A disease affects 1 in 1 000. The test finds 99 % of the sick (the sensitivity) and gives a false positive in 5 % of cases among the healthy.
You test positive. How likely is it that you are ill?
Think in numbers of people instead of probabilities — that makes it almost obvious. Out of 100 000 people:
| Number | Positives | |
|---|---|---|
| Ill | 100 | 99 |
| Healthy | 99 900 | 4 995 |
| Total positives | 5 094 |
Most people guess 95–99 %. The right answer is under 2 %.
Why? There are so enormously many more healthy people that even a small false-positive rate gives more false alarms than genuine ones. Ignoring how rare the disease is is called the base rate fallacy, and it is hard to stop doing even when you know about it.
This is not a curiosity. The same mathematics holds for every rare event:
| Application | The rare class |
|---|---|
| Fraud detection | ~0.1 % of the transactions |
| Face recognition in a crowd | the wanted person |
| Screening for rare diseases | the ill |
| Alerts in a monitoring system | real incidents |
A system with 99 % accuracy on a class that makes up 0.1 % of the cases produces an overwhelming number of false alarms. That is why precision and recall are reported separately, and why «accuracy» is an almost meaningless measure with imbalanced classes.
Bayes in sequential form. The posterior becomes the next prior:
That is exactly how AI-grafen's mastery model works: every answered exercise updates the belief that the learner has mastered the node, with the previous estimate as the starting point.
Code
def bayes(prior, sensitivity, false_positive):
"""P(ill | a positive test)."""
true_pos = sensitivity * prior
false_pos = false_positive * (1 - prior)
return true_pos / (true_pos + false_pos)
print(round(bayes(0.001, 0.99, 0.05) * 100, 1), "%") # 1.9 %
# How the prior dominates
for prior in (0.001, 0.01, 0.1, 0.5):
print(f" prior {prior:<6} → after a positive test: {bayes(prior, 0.99, 0.05):.1%}")
# prior 0.001 → 1.9%
# prior 0.01 → 16.7%
# prior 0.1 → 68.8%
# prior 0.5 → 95.2%
# Two independent positive tests: the posterior becomes the next prior
p = 0.001
for round_ in (1, 2):
p = bayes(p, 0.99, 0.05)
print(f" after {round_} positive tests: {p:.1%}")
# after 1 positive tests: 1.9%
# after 2 positive tests: 28.2%
# The same mathematics as an imbalanced classification problem
N, fraud_rate = 1_000_000, 0.001
true_pos = int(N * fraud_rate * 0.99)
false_pos = int(N * (1 - fraud_rate) * 0.05)
print(f"alerts: {true_pos + false_pos:,} of which {true_pos:,} genuine "
f"→ precision {true_pos / (true_pos + false_pos):.1%}")
# alerts: 50,940 of which 990 genuine → precision 1.9%
The last block is the same calculation as the medical test, put differently: a reviewer would have to look at almost 51 000 alerts in order to find 990 real cases — 51 alerts per genuine hit.
Mastery means
- Uses Bayes' theorem numerically
- Explains the base rate fallacy
- Applies Bayes to a diagnostic test
Sign in to do the exercises and build your mastery up.
Sources
- Matteboken (Mattecentrum) — free to read, non-profit association
- Khan Academy — matematik — CC BY-NC-SA 3.0
- Wikipedia — Bayes' theorem (CC BY-SA 4.0) — CC BY-SA 4.0