Specification problems and proxy objectives
Be able to give examples of when what you measure differs from what you want, and to design better objectives.
Prerequisites
- FReward hackingrequired
Intuition
A specification problem arises when what you can measure is not what you actually want. You optimise the proxy, and the system gets good at the proxy — not at the goal.
Goodhart's law: «when a measure becomes a target it ceases to be a good measure.»
| What you want | What you measure | What you get |
|---|---|---|
| The pupil learning | time in the app | an app that is hard to leave |
| Satisfied users | thumbs up | a model that agrees with everything |
| Good code | lines of code | diluted code |
| Safe answers | the share of questions declined | a model that refuses everything |
| Relevant search results | the click-through rate | clickbait |
It is not a technical fault but a design fault. The system does exactly what you asked for — the problem is what you asked for.
And it applies just as much to people and organisations as to RL agents. Reward hacking in a game and a sales organisation optimising its quarterly figures are the same phenomenon.
Formal
Why the proxy breaks down precisely when you push it. A proxy correlates with the goal within the region you normally inhabit. Optimise hard on and you end up outside that region, and there the correlation no longer holds.
Goodhart's law is therefore not mysterious: it is extrapolation beyond the measured range, exactly as when you extend a regression line too far.
Four forms, in Manheim and Garrabrant's taxonomy:
| Form | The mechanism |
|---|---|
| Regressional | the proxy contains noise; picking the maximum also picks the maximum noise |
| Extremal | the relationship breaks down in the extremes |
| Causal | you optimise a correlation with no causal link |
| Adversarial | somebody actively exploits the proxy |
Countermeasures:
| Countermeasure | The idea |
|---|---|
| Several objectives at once | harder to game all of them at the same time |
| Guard rails | metrics that must not get worse, even if the main metric rises |
| Satisficing | reach «good enough» rather than maximising |
| Regularisation against a reference | a KL penalty against the starting model in RLHF |
| Human review | of samples, particularly in the tails |
| Rotate the metrics | makes it harder to optimise against a fixed target |
The KL penalty in RLHF is the clearest example of a countermeasure in practice: the reward model is a proxy for human preferences, and without the penalty the policy drifts off into text that the reward model loves but humans do not.
The question to always ask: «how would I maximise this metric if I did not care at all about the underlying purpose?» If you have an answer in under a minute, the metric is too gameable.
For AI-grafen concretely. An obvious proxy objective would be «the share of correct answers». That is maximised by making the exercises easier — which is the opposite of the purpose. So what is measured instead is mastered nodes per study hour spent, with the share of tutor answers that give away the answer key as a guard rail, and the share of transfer exercises passed as a separate metric: those cannot be passed by recognition.
Code
import numpy as np
from dataclasses import dataclass
@dataclass
class Metric:
name: str
role: str # north_star | support | guard_rail
direction: str # up | down | stable
threshold: float | None = None
def gameability_analysis(metric: str, ways_to_game: list[str]):
"""Document how the metric can be maximised without the purpose being met."""
return {"metric": metric, "ways": ways_to_game, "count": len(ways_to_game),
"verdict": "too gameable" if len(ways_to_game) >= 3 else "acceptable"}
print(gameability_analysis("the share of correct answers", [
"make the exercises easier",
"give away the answer key in the tutor text",
"let the pupil try an unlimited number of times",
"remove the hard nodes from the path",
]))
# Regressional Goodhart: picking the maximum also picks the maximum noise
rng = np.random.default_rng(0)
n = 10_000
true_value = rng.normal(0, 1, n)
noise = rng.normal(0, 1, n)
proxy = true_value + noise # a correlation of ~0.71
for top in (1.0, 0.1, 0.01, 0.001):
k = max(1, int(n * top))
chosen = np.argsort(-proxy)[:k]
print(f" top {top:>6.1%}: proxy {proxy[chosen].mean():>6.3f} "
f"true value {true_value[chosen].mean():>6.3f} "
f"gap {proxy[chosen].mean() - true_value[chosen].mean():>5.3f}")
# ↑ the harder the selection, the larger the share of the proxy that is pure noise
# Guard rails: block a release that makes them worse, whatever the main metric does
METRICS = [
Metric("mastered_nodes_per_hour", "north_star", "up"),
Metric("share_of_transfer_exercises_passed", "support", "up"),
Metric("share_of_answers_giving_away_the_key", "guard_rail", "down", threshold=0.02),
Metric("share_abandoning_mid_path", "guard_rail", "down", threshold=0.15),
]
def review_release(before: dict, after: dict):
rows, blocks = [], False
for m in METRICS:
d = after[m.name] - before[m.name]
breach = m.threshold is not None and after[m.name] > m.threshold
if breach and m.role == "guard_rail":
blocks = True
rows.append({"metric": m.name, "delta": round(d, 4), "breach": breach})
return {"blocks": blocks, "rows": rows}
# A KL penalty against the reference model — the countermeasure in RLHF
def reward_with_kl(reward, log_p_policy, log_p_reference, beta=0.05):
return reward - beta * (log_p_policy - log_p_reference)
# Without the KL term the policy drifts towards text the reward model loves
# but humans do not — extremal Goodhart in its purest form.
The table in the middle is Goodhart's law in numbers: at the top 0.1 % the gap between the proxy and the true value is nearly as large as the whole true value.
Mastery means
- Recognises proxy objectives and their weaknesses
- Anticipates how an objective can be gamed
- Designs objectives with guard rails
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Categorizing Variants of Goodhart's Law — arXiv (open access; licence per article)
- arXiv — Concrete Problems in AI Safety — arXiv (open access; licence per article)
- arXiv — Defining and Characterizing Reward Hacking — arXiv (open access; licence per article)