Reward hacking
Be able to give examples of reward hacking and design rewards that are harder to game.
Prerequisites
Intuition
Reward hacking is when a system maximises the measure but not the goal. It is not an AI problem but a measurement problem — Goodhart's law: when a measure becomes a target it ceases to be a good measure.
Classic examples from RL:
- A boat in a racing game that discovers it can spin in a circle and collect point items forever instead of finishing the race.
- A robot arm meant to grasp an object learns to put its hand between the camera and the object — it looks like a grasp.
In language models:
- Answers that are long, structured and confident are rewarded by raters — regardless of correctness.
- The model agrees with the user (sycophancy) because agreement was chosen more often in the annotation.
- «I'm afraid I can't help with that» is a safe answer that never scores low on harmfulness.
Formal
Why it arises systematically: the reward is always a proxy for the true objective . They correlate in the training distribution. When the policy is optimised hard it ends up in the tail of the distribution — and there the correlation stops. Gao et al. (2022) measured this: the true reward rises, reaches a peak and then falls while the proxy reward keeps going up. The phenomenon scales predictably with the KL distance from the reference policy.
Countermeasures:
| Measure | Mechanism |
|---|---|
| A KL penalty | stops the policy reaching the tail where the proxy breaks down |
| A reward ensemble | several reward models; hacking one is easier than hacking all |
| Early stopping on the true metric | measure against human judgement, not the proxy |
| Length normalisation | removes the most common shortcut |
| Adversarial data collection | annotate exactly the cases where the model has started to drift |
| Several metrics in tension | helpfulness and correctness and concision |
The practical test: have humans rate a sample at several points during training. If the proxy reward diverges from the human judgement, the hacking has begun — stop there, not when the proxy levels off.
Code
# Diagnostics during preference training: follow the proxy and the true metric in parallel
def training_diagnostics(step, policy, reward_model, human_sample, ref):
answers = [policy(p) for p in EVAL_PROMPTS]
return {
"step": step,
"proxy_reward": float(np.mean([reward_model(p, s) for p, s in zip(EVAL_PROMPTS, answers)])),
"human": human_sample(answers), # expensive — run every nth step
"mean_length": float(np.mean([len(s.split()) for s in answers])),
"kl_vs_ref": kl_divergence(policy, ref, EVAL_PROMPTS),
"agreement_share": float(np.mean([starts_with_agreement(s) for s in answers])),
"refusal_share": float(np.mean([refuses(s) for s in answers])),
}
# step proxy human length KL agreement refusal
# 0 0.51 0.62 118 0.0 0.12 0.03
# 500 0.68 0.71 142 2.1 0.15 0.04
# 1000 0.79 0.73 198 5.8 0.24 0.06
# 1500 0.86 0.69 287 11.4 0.38 0.11 ← proxy up, human DOWN: stop at ~1000
The table is the whole point: without the «human» column you would have trained on and believed the model was getting better.
Mastery means
- Gives concrete examples of reward hacking
- Designs rewards that are harder to game
- Detects hacking before it reaches production
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Scaling Laws for Reward Model Overoptimization — arXiv (open access; licence per article)
- arXiv — Concrete Problems in AI Safety — arXiv (open access; licence per article)
- arXiv — Language Models Learn to Mislead Humans via RLHF — arXiv (open access; licence per article)