Skip to content
AI-grafen
FAI engineeringReinforcement learning· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Actor–critic and PPO

Be able to explain the advantage, the clipping and why PPO is used in RLHF.

Prerequisites

Intuition

Actor–critic combines the two families:

  • The actor — the policy πθ(a∣s)\pi_\theta(a\mid s), the one that acts.
  • The critic — the value function Vϕ(s)V_\phi(s), the one that judges how good the situation is.

The critic is used as a baseline, so that the actor is told not just «that went well» but «that went better than expected». That is the advantage:

At=Gt−Vϕ(st)A_t = G_t - V_\phi(s_t)

The difference is large in practice. «You got 100 points» says little if the situation always gives 100. «You got 20 more than expected» is a usable signal.

PPO's contribution is to solve a different problem: how large a step is allowed? Too large a policy step can destroy a working policy, and unlike supervised learning it does not come back — the next round of data is collected by the broken policy.

PPO's answer is to clip away the incentive to move too far. No hard ceiling, no cluster of KL computations — just a min with a clipped alternative.

Formal

GAE (generalized advantage estimation) interpolates between low bias and low variance:

δt=rt+γV(st+1)−V(st),AtGAE=∑l≥0(γλ)lδt+l\delta_t = r_t + \gamma V(s_{t+1}) - V(s_t), \qquad A_t^{\text{GAE}} = \sum_{l \geq 0} (\gamma\lambda)^l \delta_{t+l}

λ=0\lambda = 0 gives pure TD (low variance, high bias), λ=1\lambda = 1 gives Monte Carlo (high variance, low bias). Typically λ=0.95\lambda = 0.95.

PPO's clipped objective. With the ratio rt(θ)=πθ(at∣st)πθold(at∣st)r_t(\theta) = \dfrac{\pi_\theta(a_t\mid s_t)}{\pi_{\theta_{\text{old}}}(a_t\mid s_t)}:

LCLIP=Et[min⁡(rtAt,  clip(rt,1−ϵ,1+ϵ) At)]L^{\text{CLIP}} = \mathbb{E}_t\left[\min\left(r_t A_t,\; \mathrm{clip}(r_t, 1-\epsilon, 1+\epsilon)\, A_t\right)\right]

The asymmetry is the whole point, and it is often missed. With ϵ=0.2\epsilon = 0.2:

AtA_trtr_tThe clipped termmin picksThe effect
+11.51.21.2the upside is capped — stop increasing the probability
+10.50.80.5no capping — free to recover
−11.5−1.2−1.5no capping — free to be punished in full
−10.5−0.8−0.8the downside is capped — stop decreasing the probability

In other words: min only clips in the direction that would make the update larger. Moving back towards the old policy is always allowed.

The full loss has three terms:

L=−LCLIP+c1(Vϕ−G)2⏟critic−c2H[πθ]⏟entropyL = -L^{\text{CLIP}} + c_1 \underbrace{(V_\phi - G)^2}_{\text{critic}} - c_2 \underbrace{H[\pi_\theta]}_{\text{entropy}}

The entropy bonus counteracts the policy collapsing to determinism too early.

Why PPO in RLHF? Four reasons, in order:

  1. Robust to hyperparameters — important when every run costs GPU hours and you cannot search broadly.
  2. Several epochs per data collection — the ratio makes it safe to reuse data, which is decisive when the «environment» is expensive human or modelled feedback.
  3. No second-order machinery — TRPO needs conjugate gradients and Fisher matrix products; PPO is a few lines.
  4. The KL penalty against the reference model fits in naturally as an extra term — you want the model to improve without drifting away from the language model it started from.

DPO and other direct methods have in recent years taken over much of the RLHF work precisely because they avoid the whole RL loop — but PPO is still the reference implementation to understand first.

Code

import torch

def gae(rewards, values, dones, gamma=0.99, lam=0.95):
    """values has length T+1 (the last one is V(s_T))."""
    T = len(rewards)
    A = torch.zeros(T)
    last = 0.0
    for t in reversed(range(T)):
        non_terminal = 1.0 - dones[t]
        delta = rewards[t] + gamma * values[t + 1] * non_terminal - values[t]
        last = delta + gamma * lam * non_terminal * last
        A[t] = last
    return A, A + values[:T]                     # (advantages, value targets)

def ppo_update(policy, critic, opt, batch, epochs=4, eps=0.2, c1=0.5, c2=0.01):
    s, a, old_logp, A, V_target = batch
    A = (A - A.mean()) / (A.std() + 1e-8)
    for _ in range(epochs):                       # several epochs on the SAME data — what PPO enables
        d = policy(s)
        logp = d.log_prob(a)
        ratio = torch.exp(logp - old_logp)
        unclipped = ratio * A
        clipped = torch.clamp(ratio, 1 - eps, 1 + eps) * A
        pol_loss = -torch.min(unclipped, clipped).mean()
        v_loss = ((critic(s).squeeze(-1) - V_target) ** 2).mean()
        entropy = d.entropy().mean()
        loss = pol_loss + c1 * v_loss - c2 * entropy
        opt.zero_grad(); loss.backward()
        torch.nn.utils.clip_grad_norm_(
            list(policy.parameters()) + list(critic.parameters()), 0.5)
        opt.step()

        if (logp - old_logp).pow(2).mean() > 0.03:   # approximate KL — break if we have drifted too far
            break

Three diagnostic quantities to log, and what they mean:

QuantityA healthy levelWhat a deviation means
approximate KL per update0.005–0.02too high: too many epochs or too high an lr
the share of clipped samples0.1–0.3near 0: the clipping does nothing; near 1: the steps are too large
the policy's entropyfalls slowlya steep fall: collapse, raise c2

The combination «a low clip share and a high KL» almost always means the value function is broken rather than the policy.

Mastery means

  • Explains actor–critic and advantage estimation
  • Derives and interprets PPO's clipped objective
  • Knows why PPO was chosen for RLHF

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences