Skip to content
AI-grafen
CBuilderStatistics and probability· about 30 min· fundamentals that rarely change· verified 2026-09-21· EN

Sampling: who did we ask?

Be able to explain why a biased sample gives wrong answers — and link this to biased training data.

Prerequisites

Everyday explanation

You want to know what all employees in the company think about the new wellness agreement. You cannot ask everyone — so you ask a few. This is called a sample.

But who you ask determines the answer.

If you askYou find out
Only those who go to the gymwhat those who already train think — not those who have no time
Only your teamwhat your team thinks
Only those who reply to the emailwhat those who read emails often think
Only on Fridayswhat people think about Friday activities

A sample that does not reflect the whole is called biased. The answer becomes wrong — not because someone lied, but because the wrong people were asked.

A good sample: randomly selected people from the entire group. Draw names from a hat. Then all types are represented roughly as much as they exist.

Intuition

Three ways a sample becomes biased:

ErrorWhat happens
Wrong group askedyou ask those who are easy to reach
Not everyone respondsthose who respond differ from those who do not
Survivorship biasyou only see those who remain

The third is the tricky one. A classic example: during World War II, researchers examined where planes had the most bullet holes to reinforce armour there. But they could only examine planes that came back. The bullet holes were therefore in places where a plane could be hit and still survive — and the armour should have been placed where there were no holes.

The same applies to AI. An AI learns from data — and the data is a sample.

If the AI was trained onIt works worse for
Mostly English textSwedish, and even worse for minority languages
Mostly light-skinned facesdark-skinned faces
Mostly male voicesfemale and children’s voices
Data from a big citythe countryside

This is not hypothetical. Early facial recognition systems had much higher error rates for dark-skinned women than for light-skinned men — because the training data was biased.

The question to always ask about an AI system: who was in the data, and who was not?

Interactive

Task 1 — find the bias.

For each survey: who is missed, and how does it affect the answer?

#SurveyWho is missed?
1“How often do you cycle to work?” — asked at the bike rack
2“What do you think of the app?” — survey in the app
3“How are you doing?” — asked in the conference room during a meeting
4“Is public transport good?” — asked on the bus
5“How satisfied are you?” — email survey to all customers

Answer key: 1 those who do not cycle · 2 those who stopped using the app (and they are the most dissatisfied) · 3 those who are absent — often those who are feeling worst or are very busy · 4 those who stopped riding · 5 those who do not have the energy to answer, which is often the indifferent and the busiest.

Numbers 2 and 3 are the most instructive: in both cases those who would have given the most critical answer are systematically missing.

Task 2 — make a good sample.

You need to find out what all employees think about the wellness agreement. Write down:

  • How do you choose who to ask?
  • How many do you need to ask?
  • What do you do with those who do not want to answer?
  • How do you check that your sample resembles the whole company? (For example, compare the proportion per department.)

Task 3 — link to AI.

Imagine an AI that is supposed to recognise Swedish words in speech. If it was mostly trained on adults from Stockholm — for whom will it work worse? Write three groups and justify.

Mastery means

  • Explains what a representative sample is
  • Recognises biased samples
  • Links to biased training data in AI

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences