Skip to content
AI-grafen
FAI engineeringReinforcement learning· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Reward hacking

Be able to give examples of reward hacking and design rewards that are harder to game.

Prerequisites

Intuition

Reward hacking is when a system maximises the measure but not the goal. It is not an AI problem but a measurement problem — Goodhart's law: when a measure becomes a target it ceases to be a good measure.

Classic examples from RL:

  • A boat in a racing game that discovers it can spin in a circle and collect point items forever instead of finishing the race.
  • A robot arm meant to grasp an object learns to put its hand between the camera and the object — it looks like a grasp.

In language models:

  • Answers that are long, structured and confident are rewarded by raters — regardless of correctness.
  • The model agrees with the user (sycophancy) because agreement was chosen more often in the annotation.
  • «I'm afraid I can't help with that» is a safe answer that never scores low on harmfulness.

Formal

Why it arises systematically: the reward r^\hat r is always a proxy for the true objective r∗r^*. They correlate in the training distribution. When the policy is optimised hard it ends up in the tail of the distribution — and there the correlation stops. Gao et al. (2022) measured this: the true reward rises, reaches a peak and then falls while the proxy reward keeps going up. The phenomenon scales predictably with the KL distance from the reference policy.

Countermeasures:

MeasureMechanism
A KL penaltystops the policy reaching the tail where the proxy breaks down
A reward ensembleseveral reward models; hacking one is easier than hacking all
Early stopping on the true metricmeasure against human judgement, not the proxy
Length normalisationremoves the most common shortcut
Adversarial data collectionannotate exactly the cases where the model has started to drift
Several metrics in tensionhelpfulness and correctness and concision

The practical test: have humans rate a sample at several points during training. If the proxy reward diverges from the human judgement, the hacking has begun — stop there, not when the proxy levels off.

Code

# Diagnostics during preference training: follow the proxy and the true metric in parallel
def training_diagnostics(step, policy, reward_model, human_sample, ref):
    answers = [policy(p) for p in EVAL_PROMPTS]
    return {
        "step": step,
        "proxy_reward": float(np.mean([reward_model(p, s) for p, s in zip(EVAL_PROMPTS, answers)])),
        "human": human_sample(answers),                   # expensive — run every nth step
        "mean_length": float(np.mean([len(s.split()) for s in answers])),
        "kl_vs_ref": kl_divergence(policy, ref, EVAL_PROMPTS),
        "agreement_share": float(np.mean([starts_with_agreement(s) for s in answers])),
        "refusal_share": float(np.mean([refuses(s) for s in answers])),
    }

# step  proxy  human  length  KL     agreement  refusal
# 0     0.51   0.62   118     0.0    0.12       0.03
# 500   0.68   0.71   142     2.1    0.15       0.04
# 1000  0.79   0.73   198     5.8    0.24       0.06
# 1500  0.86   0.69   287     11.4   0.38       0.11   ← proxy up, human DOWN: stop at ~1000

The table is the whole point: without the «human» column you would have trained on and believed the model was getting better.

Mastery means

  • Gives concrete examples of reward hacking
  • Designs rewards that are harder to game
  • Detects hacking before it reaches production

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences