Experiment: biased training data and its consequences
Be able to identify and handle bias in training data, and demonstrate how the model inherits these patterns.
Prerequisites
- BTraining data, features and labelsrequired
- CManual data labellingrequired
Intuition
The model learns the patterns present in the data — including those you did not intend.
If you train a classifier on 'defective/okay' with 90 images of defective items on a white table and 90 images of okay items on a dark mat, the model learns 'white surface = defective'. A defective item on the mat is classified as 'okay'.
This is called bias in the data. The model is not logically flawed — it found the simplest correlating pattern that separated the categories. The solution is almost always better data, not a more complex model.
Interactive
Experiment in Teachable Machine:
- Class A: 20 images of invoices — all scanned against a white background.
- Class B: 20 images of receipts — all photographed against a dark background.
- Train. Test with an invoice on a dark background. What does the model say? Often 'receipt'.
- Repeat with mixed backgrounds for both classes. Test again.
Write down: accuracy on standard images, accuracy on 'swapped background', before and after. This is a controlled experiment: only the data changed.
Mastery means
- Train a model on biased data and demonstrate that it inherits the bias
- Propose a fix in the data, not in the model
Sign in to do the exercises and build your mastery up.
Sources
- Teachable Machine (Google, gratis) — free web service
- Wikipedia — Algoritmisk partiskhet (CC BY-SA 4.0) — CC BY-SA 4.0