Skip to content
AI-grafen
FAI engineeringAI safety and alignment· about 90 min· fast-moving, sources checked often· verified 2026-09-21· EN

Specification problems and proxy objectives

Be able to give examples of when what you measure differs from what you want, and to design better objectives.

Prerequisites

Intuition

A specification problem arises when what you can measure is not what you actually want. You optimise the proxy, and the system gets good at the proxy — not at the goal.

Goodhart's law: «when a measure becomes a target it ceases to be a good measure.»

What you wantWhat you measureWhat you get
The pupil learningtime in the appan app that is hard to leave
Satisfied usersthumbs upa model that agrees with everything
Good codelines of codediluted code
Safe answersthe share of questions declineda model that refuses everything
Relevant search resultsthe click-through rateclickbait

It is not a technical fault but a design fault. The system does exactly what you asked for — the problem is what you asked for.

And it applies just as much to people and organisations as to RL agents. Reward hacking in a game and a sales organisation optimising its quarterly figures are the same phenomenon.

Formal

Why the proxy breaks down precisely when you push it. A proxy PP correlates with the goal MM within the region you normally inhabit. Optimise hard on PP and you end up outside that region, and there the correlation no longer holds.

Goodhart's law is therefore not mysterious: it is extrapolation beyond the measured range, exactly as when you extend a regression line too far.

Four forms, in Manheim and Garrabrant's taxonomy:

FormThe mechanism
Regressionalthe proxy contains noise; picking the maximum also picks the maximum noise
Extremalthe relationship breaks down in the extremes
Causalyou optimise a correlation with no causal link
Adversarialsomebody actively exploits the proxy

Countermeasures:

CountermeasureThe idea
Several objectives at onceharder to game all of them at the same time
Guard railsmetrics that must not get worse, even if the main metric rises
Satisficingreach «good enough» rather than maximising
Regularisation against a referencea KL penalty against the starting model in RLHF
Human reviewof samples, particularly in the tails
Rotate the metricsmakes it harder to optimise against a fixed target

The KL penalty in RLHF is the clearest example of a countermeasure in practice: the reward model is a proxy for human preferences, and without the penalty the policy drifts off into text that the reward model loves but humans do not.

The question to always ask: «how would I maximise this metric if I did not care at all about the underlying purpose?» If you have an answer in under a minute, the metric is too gameable.

For AI-grafen concretely. An obvious proxy objective would be «the share of correct answers». That is maximised by making the exercises easier — which is the opposite of the purpose. So what is measured instead is mastered nodes per study hour spent, with the share of tutor answers that give away the answer key as a guard rail, and the share of transfer exercises passed as a separate metric: those cannot be passed by recognition.

Code

import numpy as np
from dataclasses import dataclass

@dataclass
class Metric:
    name: str
    role: str                   # north_star | support | guard_rail
    direction: str              # up | down | stable
    threshold: float | None = None

def gameability_analysis(metric: str, ways_to_game: list[str]):
    """Document how the metric can be maximised without the purpose being met."""
    return {"metric": metric, "ways": ways_to_game, "count": len(ways_to_game),
            "verdict": "too gameable" if len(ways_to_game) >= 3 else "acceptable"}

print(gameability_analysis("the share of correct answers", [
    "make the exercises easier",
    "give away the answer key in the tutor text",
    "let the pupil try an unlimited number of times",
    "remove the hard nodes from the path",
]))

# Regressional Goodhart: picking the maximum also picks the maximum noise
rng = np.random.default_rng(0)
n = 10_000
true_value = rng.normal(0, 1, n)
noise = rng.normal(0, 1, n)
proxy = true_value + noise                     # a correlation of ~0.71

for top in (1.0, 0.1, 0.01, 0.001):
    k = max(1, int(n * top))
    chosen = np.argsort(-proxy)[:k]
    print(f"  top {top:>6.1%}: proxy {proxy[chosen].mean():>6.3f}  "
          f"true value {true_value[chosen].mean():>6.3f}  "
          f"gap {proxy[chosen].mean() - true_value[chosen].mean():>5.3f}")
#  ↑ the harder the selection, the larger the share of the proxy that is pure noise

# Guard rails: block a release that makes them worse, whatever the main metric does
METRICS = [
    Metric("mastered_nodes_per_hour", "north_star", "up"),
    Metric("share_of_transfer_exercises_passed", "support", "up"),
    Metric("share_of_answers_giving_away_the_key", "guard_rail", "down", threshold=0.02),
    Metric("share_abandoning_mid_path", "guard_rail", "down", threshold=0.15),
]

def review_release(before: dict, after: dict):
    rows, blocks = [], False
    for m in METRICS:
        d = after[m.name] - before[m.name]
        breach = m.threshold is not None and after[m.name] > m.threshold
        if breach and m.role == "guard_rail":
            blocks = True
        rows.append({"metric": m.name, "delta": round(d, 4), "breach": breach})
    return {"blocks": blocks, "rows": rows}

# A KL penalty against the reference model — the countermeasure in RLHF
def reward_with_kl(reward, log_p_policy, log_p_reference, beta=0.05):
    return reward - beta * (log_p_policy - log_p_reference)
# Without the KL term the policy drifts towards text the reward model loves
# but humans do not — extremal Goodhart in its purest form.

The table in the middle is Goodhart's law in numbers: at the top 0.1 % the gap between the proxy and the true value is nearly as large as the whole true value.

Mastery means

  • Recognises proxy objectives and their weaknesses
  • Anticipates how an objective can be gamed
  • Designs objectives with guard rails

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences