Skip to content
AI-grafen
FAI engineeringAI safety and alignment· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

RLHF and preference learning

Be able to explain the reward model, PPO/DPO and why preference data shapes the model's behaviour.

Prerequisites

Intuition

After instruction fine-tuning (SFT) the model can answer. Preference learning teaches it how the answers should be.

Classic RLHF in three steps:

  1. Collect preferences. People are given two answers to the same prompt and choose the better one.
  2. Train a reward model r(prompt, answer) that predicts which answer people prefer.
  3. Optimise the policy (the language model) against the reward with RL — plus a KL penalty against the reference model.

The KL penalty is not a detail but what holds the whole thing together: without it the policy quickly finds texts that fool the reward model but are incomprehensible to humans.

Formal

The RLHF objective: max⁡π Ex∼D, y∼π(⋅∣x)[rϕ(x,y)]−β DKL(π(⋅∣x) ∥ πref(⋅∣x))\max_\pi\ \mathbb E_{x\sim D,\, y\sim\pi(\cdot|x)}\big[r_\phi(x,y)\big] - \beta\, D_{KL}\big(\pi(\cdot|x)\,\|\,\pi_{\text{ref}}(\cdot|x)\big)

The reward model is trained with Bradley–Terry on preference pairs: Lr=−E(x,yw,yl)[log⁡σ(rϕ(x,yw)−rϕ(x,yl))]\mathcal L_r = -\mathbb E_{(x,y_w,y_l)}\big[\log \sigma\big(r_\phi(x,y_w) - r_\phi(x,y_l)\big)\big]

DPO (Rafailov et al. 2023) shows that the optimisation problem has a closed-form solution that makes the reward model and the RL loop unnecessary — you can train directly on the preference pairs: LDPO=−E[log⁡σ(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))]\mathcal L_{DPO} = -\mathbb E\Big[\log\sigma\Big(\beta\log\frac{\pi_\theta(y_w|x)}{\pi_{ref}(y_w|x)} - \beta\log\frac{\pi_\theta(y_l|x)}{\pi_{ref}(y_l|x)}\Big)\Big]

DPO is simpler (no reward model, no PPO), more stable and cheaper — and therefore the most common choice in open models today. PPO can still win when you have a lot of preference data and want to iterate on the reward model.

Known side effects:

  • Reward hacking: the policy finds patterns the reward model rewards but humans do not like (excessive politeness, hedging, lists).
  • Length bias: longer answers are preferred in the data → the model becomes verbose. The countermeasure: a length-normalised reward.
  • Sycophancy: the model agrees with the user, because agreement was often preferred in the annotation.

Code

import torch, torch.nn.functional as F

def dpo_loss(policy_logps_w, policy_logps_l, ref_logps_w, ref_logps_l, beta=0.1):
    """logps = the summed log probability of the answer given the prompt. w = chosen, l = rejected."""
    policy_diff = policy_logps_w - policy_logps_l
    ref_diff = ref_logps_w - ref_logps_l
    return -F.logsigmoid(beta * (policy_diff - ref_diff)).mean()

# Diagnostics that reveal reward hacking early:
#  - the mean answer length per training step (rising sharply? length bias)
#  - the KL against the reference model (running away? the policy is drifting)
#  - a capability suite before and after (is the model losing general ability?)
#  - the share of answers beginning with agreement (sycophancy)

Choosing beta: a small β = more freedom to optimise the reward but a greater risk of hacking; a large β = safer but less effect. 0.1 is a common starting point, and the value should be justified by measurement, not by habit.

Mastery means

  • Explains the reward model, PPO and DPO
  • Justifies the KL penalty
  • Recognises reward hacking and sycophancy

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences