Manual data labelling
Be able to label 30 examples, identify ambiguous cases, and understand that labelling is a deliberate choice.
Prerequisites
- BTraining data, features and labelsrequired
Intuition
A model learns from labelled examples: an image with the label “cat”, a customer review with the label “positive”. Someone must set these labels manually. It sounds simple until you do it yourself.
“Customer service was not bad.” — Positive? “Good delivery, but the product was broken.” — ?
Every ambiguous case is a choice. Two people can label the same text differently. That choice is encoded into the model. That is why we write instructions (guidelines) and measure how often the labelers agree with each other.
Interactive
Do this: Write 30 short messages (or extract them from a customer chat you have access to). Label each message as “friendly”, “neutral”, or “rude”. Have a colleague label the same 30 messages without seeing your answers.
Calculate: How many times did you agree on the label? 24 out of 30 = 80% agreement. Review the 6 where you disagreed. Write a rule that resolves them (e.g. “Irony counts as rude if …”). Relabel. Did the agreement increase?
This is exactly how professional datasets are built.
Mastery means
- Labels examples according to instructions and identifies ambiguous cases
- Explains that labelling is a choice that affects model behaviour
Sign in to do the exercises and build your mastery up.
Sources
- CS Unplugged (CC BY-SA 4.0) — CC BY-SA 4.0
- Wikipedia — Inter-rater reliability (CC BY-SA 4.0) — CC BY-SA 4.0