Calibration and uncertainty in models
Be able to measure calibration and make a model express uncertainty.
Prerequisites
- CProbability — the basicsrequiredPractise in Mattegrafen ↗
- EModel evaluationrequired
Intuition
A calibrated model that says «80 % sure» is right in 80 % of those cases. Most modern networks are overconfident: they say 95 % and are right 80 % of the time.
It matters as soon as you want to act on the confidence: escalate uncertain cases to a person, abstain from answering, or sort by risk. An uncalibrated confidence makes every such rule worthless.
Measuring it: divide the predictions into ten confidence bins, compare the average confidence with the actual accuracy in each. The difference, weighted by the count, is the ECE (expected calibration error).
Code
import numpy as np
def ece(confidence, correct, bins=10):
confidence, correct = np.asarray(confidence), np.asarray(correct, float)
edges = np.linspace(0, 1, bins + 1)
error, rows = 0.0, []
for lo, hi in zip(edges[:-1], edges[1:]):
m = (confidence > lo) & (confidence <= hi)
if m.sum() == 0:
continue
acc, conf = correct[m].mean(), confidence[m].mean()
error += m.mean() * abs(acc - conf)
rows.append((f"{lo:.1f}-{hi:.1f}", int(m.sum()), round(conf, 3), round(acc, 3)))
return error, rows
def temperature_scale(logits, y_val, temperatures=np.linspace(0.5, 3.0, 26)):
"""Find the T that minimises the NLL on validation data. It does not change which class is chosen."""
def nll(T):
z = logits / T
z = z - z.max(axis=1, keepdims=True)
logp = z - np.log(np.exp(z).sum(axis=1, keepdims=True))
return -logp[np.arange(len(y_val)), y_val].mean()
return float(min(temperatures, key=nll))
T = temperature_scale(val_logits, y_val) # 1.7, say → the model was overconfident
Temperature scaling is the simplest effective method: a single parameter fitted on validation data, which does not change the ranking (and therefore not the accuracy) but makes the probabilities honest.
For language models it is harder — there is no simple probability for «the answer is right». Practical proxies: self-consistency (the agreement between sampled answers), verbalised confidence («how sure are you, 0–100?», which is poorly calibrated but better than nothing), and token probabilities for the decisive answer span.
Mastery means
- Measures calibration with ECE and a reliability diagram
- Applies temperature scaling
- Makes a model express and act on uncertainty
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — On Calibration of Modern Neural Networks — arXiv (open access; licence per article)
- arXiv — Teaching Models to Express Their Uncertainty in Words — arXiv (open access; licence per article)