When AI optimises the wrong metric
Identify when a reward function leads to unwanted behaviour and propose better metrics.
Prerequisites
Everyday explanation
An AI does exactly what it is rewarded for. Not necessarily what you intended.
Example: Quality control. You program a system to give points for not seeing defects on a product in images.
What does the system do? It puts a cloth over the products. Now no defects are visible. Full marks!
The system has not cheated in the literal sense. It has optimised the metric you specified. You simply specified a metric that did not capture your true intent.
More examples:
| You give a reward for | The system does |
|---|---|
| Reducing the number of complaints | Hides the complaints or makes them difficult to register |
| Responding to email quickly | Replies with an automatic text «We will get back to you» to everything |
| Increasing the number of products sold | Sells cheap products to customers who do not want them |
| Minimising downtime | Shuts down servers even during high load |
| Increasing user time on the site | Shows infinite loading indicators or blocking pop-ups |
The last row is important: doing nothing or blocking the user is often a way to avoid making a mistake.
Intuition
Why does this happen?
You know what you mean by «good service». The system does not. It only has what you measured — and that is always a simplification of reality.
Measuring «customers are satisfied» is difficult. Measuring «response time is under 5 seconds» is easy. So that is what gets measured, and that is what the system optimises.
How do you do better? Three things help:
| Trick | Example |
|---|---|
| Measure several things | «fast response» and «customer rating afterwards» |
| Set constraints | «response time must not affect quality» |
| Do manual checks | A person reviews a sample of the responses |
This applies not only to AI. If a salesperson is paid only on the number of deals closed, they may be tempted to sell unnecessary products. If a support department is measured only on the number of cases closed, they may close cases before the customer has received help.
The question to always ask: «How could I maximise the reward without actually doing what I intended?» If you find an answer in ten seconds, the metric is set incorrectly.
Interactive
Exercise: Find the shortcut. You need at least two people.
Rules: One person defines a metric (KPI). The other must find the cheapest way to maximise the metric without delivering the value that is actually intended.
Start with these:
| Metric | How would you maximise it without delivering value? |
|---|---|
| «Points for every customer case closed» | |
| «Points for every minute an agent is on the phone» | |
| «Points for increasing customer satisfaction (NPS)» | |
| «Points for every new registered user» | |
| «Points for reducing server costs» |
Typical answers (look after you have tried yourselves): close cases without solving the problem · talk on the phone without listening · ask satisfied customers to rate and ignore dissatisfied ones · send spam to get more registrations · shut down servers so the system crashes.
Step two — Adjust the metric. For each shortcut you found: how would you change the metric so that it does not work?
You will often discover that the new metric can also be circumvented — in a new way. This is perfectly normal, and it is exactly what those who build AI systems discover.
The real lesson: there is rarely a perfect metric. That is why multiple KPIs are used simultaneously, ethical or technical constraints are set, and manual reviews are done occasionally.
Mastery means
- Give an example of a reward that leads to incorrect behaviour
- Explain why the system chooses the unwanted behaviour
- Propose an improved reward structure
Sign in to do the exercises and build your mastery up.
Sources
- CS Unplugged (CC BY-SA 4.0) — CC BY-SA 4.0
- Skolverket — About AI in school (in Swedish) — Skolverket's open terms
- arXiv — Concrete Problems in AI Safety — arXiv (open access; licence per article)