Bias and fairness in models
Be able to measure different outcomes between groups and explain where the skew comes from.
Prerequisites
Intuition
A model with 95 % accuracy overall can have 98 % for one group and 74 % for another. The average hides it.
Where the skew comes from:
| The source | Example |
|---|---|
| Historical data | earlier decisions were skewed → the model learns them |
| The sample | a group is underrepresented in the data |
| The labels | «a good employee» was defined by earlier managers |
| The features | a postcode as a proxy for ethnicity |
| The measurement | a sensor works worse on certain skin tones |
The first step is always the same: break every measure down by group and look. What you do not measure you cannot see.
Formal
Three common fairness measures:
- Demographic parity: — the same share of positive decisions in every group.
- Equal opportunity: — the same recall among those who actually satisfy the criterion.
- Calibration: among those given the score , the share with should be in every group.
The impossibility result (Kleinberg et al. 2016; Chouldechova 2017): if the base rates differ between the groups, calibration and equal error rates cannot be satisfied at the same time, except in trivial cases. There is therefore no measure that is «the fair one» — the choice is normative and has to be justified in its context.
So the practical minimum routine is: measure per group, state which measure you prioritise and why, and let a human being decide in borderline cases.
Code
import numpy as np
def per_group(y, pred, group):
for g in np.unique(group):
m = group == g
tp = ((pred[m] == 1) & (y[m] == 1)).sum(); fp = ((pred[m] == 1) & (y[m] == 0)).sum()
fn = ((pred[m] == 0) & (y[m] == 1)).sum()
print(f"{g}: n={m.sum():4d} share positive={pred[m].mean():.2f} "
f"recall={tp / max(tp + fn, 1):.2f} precision={tp / max(tp + fp, 1):.2f}")
per_group(y_test, model.predict(X_test), group_test)
# A: n= 800 share positive=0.42 recall=0.91 precision=0.88
# B: n= 120 share positive=0.19 recall=0.63 precision=0.71 ← a large difference
Mastery means
- Measures the outcome per group
- Explains where the skew arises
- Knows that fairness measures can be incompatible
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Inherent Trade-Offs in the Fair Determination of Risk Scores — arXiv (open access; licence per article)
- Wikipedia — Algorithmic bias (CC BY-SA 4.0) — CC BY-SA 4.0