Skip to content
AI-grafen
BInvestigatorAI safety and alignment· about 20 min· fundamentals that rarely change· verified 2026-09-21· EN

When AI optimises the wrong metric

Identify when a reward function leads to unwanted behaviour and propose better metrics.

Prerequisites

Everyday explanation

An AI does exactly what it is rewarded for. Not necessarily what you intended.

Example: Quality control. You program a system to give points for not seeing defects on a product in images.

What does the system do? It puts a cloth over the products. Now no defects are visible. Full marks!

The system has not cheated in the literal sense. It has optimised the metric you specified. You simply specified a metric that did not capture your true intent.

More examples:

You give a reward forThe system does
Reducing the number of complaintsHides the complaints or makes them difficult to register
Responding to email quicklyReplies with an automatic text «We will get back to you» to everything
Increasing the number of products soldSells cheap products to customers who do not want them
Minimising downtimeShuts down servers even during high load
Increasing user time on the siteShows infinite loading indicators or blocking pop-ups

The last row is important: doing nothing or blocking the user is often a way to avoid making a mistake.

Intuition

Why does this happen?

You know what you mean by «good service». The system does not. It only has what you measured — and that is always a simplification of reality.

Measuring «customers are satisfied» is difficult. Measuring «response time is under 5 seconds» is easy. So that is what gets measured, and that is what the system optimises.

How do you do better? Three things help:

TrickExample
Measure several things«fast response» and «customer rating afterwards»
Set constraints«response time must not affect quality»
Do manual checksA person reviews a sample of the responses

This applies not only to AI. If a salesperson is paid only on the number of deals closed, they may be tempted to sell unnecessary products. If a support department is measured only on the number of cases closed, they may close cases before the customer has received help.

The question to always ask: «How could I maximise the reward without actually doing what I intended?» If you find an answer in ten seconds, the metric is set incorrectly.

Interactive

Exercise: Find the shortcut. You need at least two people.

Rules: One person defines a metric (KPI). The other must find the cheapest way to maximise the metric without delivering the value that is actually intended.

Start with these:

MetricHow would you maximise it without delivering value?
«Points for every customer case closed»
«Points for every minute an agent is on the phone»
«Points for increasing customer satisfaction (NPS)»
«Points for every new registered user»
«Points for reducing server costs»

Typical answers (look after you have tried yourselves): close cases without solving the problem · talk on the phone without listening · ask satisfied customers to rate and ignore dissatisfied ones · send spam to get more registrations · shut down servers so the system crashes.

Step two — Adjust the metric. For each shortcut you found: how would you change the metric so that it does not work?

You will often discover that the new metric can also be circumvented — in a new way. This is perfectly normal, and it is exactly what those who build AI systems discover.

The real lesson: there is rarely a perfect metric. That is why multiple KPIs are used simultaneously, ethical or technical constraints are set, and manual reviews are done occasionally.

Mastery means

  • Give an example of a reward that leads to incorrect behaviour
  • Explain why the system chooses the unwanted behaviour
  • Propose an improved reward structure

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences