RLHF and preference learning
Be able to explain the reward model, PPO/DPO and why preference data shapes the model's behaviour.
Prerequisites
- EFine-tuning language modelsrequired
- EReinforcement learning — the basicsrequired
Intuition
After instruction fine-tuning (SFT) the model can answer. Preference learning teaches it how the answers should be.
Classic RLHF in three steps:
- Collect preferences. People are given two answers to the same prompt and choose the better one.
- Train a reward model r(prompt, answer) that predicts which answer people prefer.
- Optimise the policy (the language model) against the reward with RL — plus a KL penalty against the reference model.
The KL penalty is not a detail but what holds the whole thing together: without it the policy quickly finds texts that fool the reward model but are incomprehensible to humans.
Formal
The RLHF objective:
The reward model is trained with Bradley–Terry on preference pairs:
DPO (Rafailov et al. 2023) shows that the optimisation problem has a closed-form solution that makes the reward model and the RL loop unnecessary — you can train directly on the preference pairs:
DPO is simpler (no reward model, no PPO), more stable and cheaper — and therefore the most common choice in open models today. PPO can still win when you have a lot of preference data and want to iterate on the reward model.
Known side effects:
- Reward hacking: the policy finds patterns the reward model rewards but humans do not like (excessive politeness, hedging, lists).
- Length bias: longer answers are preferred in the data → the model becomes verbose. The countermeasure: a length-normalised reward.
- Sycophancy: the model agrees with the user, because agreement was often preferred in the annotation.
Code
import torch, torch.nn.functional as F
def dpo_loss(policy_logps_w, policy_logps_l, ref_logps_w, ref_logps_l, beta=0.1):
"""logps = the summed log probability of the answer given the prompt. w = chosen, l = rejected."""
policy_diff = policy_logps_w - policy_logps_l
ref_diff = ref_logps_w - ref_logps_l
return -F.logsigmoid(beta * (policy_diff - ref_diff)).mean()
# Diagnostics that reveal reward hacking early:
# - the mean answer length per training step (rising sharply? length bias)
# - the KL against the reference model (running away? the policy is drifting)
# - a capability suite before and after (is the model losing general ability?)
# - the share of answers beginning with agreement (sycophancy)
Choosing beta: a small β = more freedom to optimise the reward but a greater risk of hacking; a large β = safer but less effect. 0.1 is a common starting point, and the value should be justified by measurement, not by habit.
Mastery means
- Explains the reward model, PPO and DPO
- Justifies the KL penalty
- Recognises reward hacking and sycophancy
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Training language models to follow instructions with human feedback — arXiv (open access; licence per article)
- arXiv — Direct Preference Optimization — arXiv (open access; licence per article)