Training data, features and labels
Identify features (inputs) and the label (target) in a table, explain what a training example is, and understand why errors or biased examples lead to a poor model.
Prerequisites
- BData in everyday liferequired
Everyday explanation
To teach a computer to predict whether a customer will churn, you show it history from many customers and mark which ones actually left. Each customer with their outcome is a training example.
The data points the computer analyses — number of purchases, support tickets, customer tenure — are called features. The outcome — staying or leaving — is called the label.
Intuition
| Number of purchases | Support tickets | Customer tenure (years) | Churns? |
|---|---|---|---|
| 5 | 0 | 3 | no |
| 1 | 4 | 0.5 | yes |
| 3 | 1 | 2 | no |
The columns on the left are features; the bold one is the label. The model learns the relationship: features → label.
If almost all examples are loyal customers with many purchases, the model learns "many purchases = stays" and misses that a customer with many purchases but high stress might churn. This is biased data — the model can never be better than the examples it was given.
Mastery means
- Identifies features and label in a dataset
- Provides an example of biased training data and its consequences
Sign in to do the exercises and build your mastery up.
Sources
Leads to
Part of the goals (35)
- AI for beginners
- Data: collect, clean, document
- Fine-tune and run your own models
- AI, ethics and society
- Classical machine learning in practice
- Training neural networks for real
- Builder — collect data, train a model and test AI critically
- Seeing and hearing with AI
- An AI service in operation
- Build a RAG system you can trust
- Classical ML for real
- Train your first neural network
- Frontier Lab — an independent research project
- Evals in practice
- Responsible AI in practice
- AI safety in practice
- Language models in practice
- Build an agent you can trust
- Build an AI service that survives production
- Build a memory system for an agent
- Interpreting a language model
- Multimodal systems
- AI in production
- Train an agent with reward
- Build an NLP system end to end
- Deep reinforcement learning
- Fine-tune a model with LoRA
- Image classification with convolutional networks
- Run models more cheaply: quantisation
- Understand how generative AI works
- Reproduce a paper
- Build a transformer from scratch
- Generative models in depth
- Build a voice interface
- Statistics for experiments