Sampling: who did we ask?
Be able to explain why a biased sample gives wrong answers — and link this to biased training data.
Prerequisites
Everyday explanation
You want to know what all employees in the company think about the new wellness agreement. You cannot ask everyone — so you ask a few. This is called a sample.
But who you ask determines the answer.
| If you ask | You find out |
|---|---|
| Only those who go to the gym | what those who already train think — not those who have no time |
| Only your team | what your team thinks |
| Only those who reply to the email | what those who read emails often think |
| Only on Fridays | what people think about Friday activities |
A sample that does not reflect the whole is called biased. The answer becomes wrong — not because someone lied, but because the wrong people were asked.
A good sample: randomly selected people from the entire group. Draw names from a hat. Then all types are represented roughly as much as they exist.
Intuition
Three ways a sample becomes biased:
| Error | What happens |
|---|---|
| Wrong group asked | you ask those who are easy to reach |
| Not everyone responds | those who respond differ from those who do not |
| Survivorship bias | you only see those who remain |
The third is the tricky one. A classic example: during World War II, researchers examined where planes had the most bullet holes to reinforce armour there. But they could only examine planes that came back. The bullet holes were therefore in places where a plane could be hit and still survive — and the armour should have been placed where there were no holes.
The same applies to AI. An AI learns from data — and the data is a sample.
| If the AI was trained on | It works worse for |
|---|---|
| Mostly English text | Swedish, and even worse for minority languages |
| Mostly light-skinned faces | dark-skinned faces |
| Mostly male voices | female and children’s voices |
| Data from a big city | the countryside |
This is not hypothetical. Early facial recognition systems had much higher error rates for dark-skinned women than for light-skinned men — because the training data was biased.
The question to always ask about an AI system: who was in the data, and who was not?
Interactive
Task 1 — find the bias.
For each survey: who is missed, and how does it affect the answer?
| # | Survey | Who is missed? |
|---|---|---|
| 1 | “How often do you cycle to work?” — asked at the bike rack | |
| 2 | “What do you think of the app?” — survey in the app | |
| 3 | “How are you doing?” — asked in the conference room during a meeting | |
| 4 | “Is public transport good?” — asked on the bus | |
| 5 | “How satisfied are you?” — email survey to all customers |
Answer key: 1 those who do not cycle · 2 those who stopped using the app (and they are the most dissatisfied) · 3 those who are absent — often those who are feeling worst or are very busy · 4 those who stopped riding · 5 those who do not have the energy to answer, which is often the indifferent and the busiest.
Numbers 2 and 3 are the most instructive: in both cases those who would have given the most critical answer are systematically missing.
Task 2 — make a good sample.
You need to find out what all employees think about the wellness agreement. Write down:
- How do you choose who to ask?
- How many do you need to ask?
- What do you do with those who do not want to answer?
- How do you check that your sample resembles the whole company? (For example, compare the proportion per department.)
Task 3 — link to AI.
Imagine an AI that is supposed to recognise Swedish words in speech. If it was mostly trained on adults from Stockholm — for whom will it work worse? Write three groups and justify.
Mastery means
- Explains what a representative sample is
- Recognises biased samples
- Links to biased training data in AI
Sign in to do the exercises and build your mastery up.
Sources
- Statistics Sweden — myndighetsmaterial
- CS Unplugged (CC BY-SA 4.0) — CC BY-SA 4.0
- Buolamwini & Gebru — Gender Shades — PMLR open access