Skip to content
AI-grafen
FAI engineeringReinforcement learning· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Policy gradient and REINFORCE

Be able to derive the policy gradient theorem and implement REINFORCE.

Prerequisites

Intuition

Q-learning learns values and derives the policy from them. Policy gradient skips the middle step and optimises the policy directly.

The policy πθ(a∣s)\pi_\theta(a \mid s) is a neural network that gives probabilities over actions. We want to adjust θ\theta so that good outcomes become more likely.

The idea in one sentence:

It went well — do more of what you did. It went badly — do less.

Mathematically: multiply the log probability of every action by how well the episode turned out, and walk up the gradient.

Why care, when Q-learning exists?

Value-basedPolicy gradient
Continuous actionshard (a max over infinitely many)natural
A stochastic policynoyes — needed in games with bluffing
Convergencecan oscillate with function approximationmore stable guarantees
Sample efficiencybetter (off-policy, replay)worse (on-policy)

The last row is the price: REINFORCE throws away all its data after every update.

Derivation

The objective is the expected return over trajectories τ\tau:

J(θ)=Eτ∼πθ[R(τ)]=∫pθ(τ)R(τ) dτJ(\theta) = \mathbb{E}_{\tau \sim \pi_\theta}[R(\tau)] = \int p_\theta(\tau) R(\tau)\, d\tau

The problem: the gradient lands on the probability distribution, which we cannot differentiate through directly.

The log-derivative trick solves it. Since ∇θpθ=pθ∇θlog⁡pθ\nabla_\theta p_\theta = p_\theta \nabla_\theta \log p_\theta:

∇θJ=∫∇θpθ(τ)R(τ) dτ=∫pθ(τ) ∇θlog⁡pθ(τ) R(τ) dτ=Eτ ⁣[∇θlog⁡pθ(τ) R(τ)]\nabla_\theta J = \int \nabla_\theta p_\theta(\tau) R(\tau)\, d\tau = \int p_\theta(\tau)\, \nabla_\theta \log p_\theta(\tau)\, R(\tau)\, d\tau = \mathbb{E}_\tau\!\left[\nabla_\theta \log p_\theta(\tau)\, R(\tau)\right]

Now it is an expectation again — and expectations can be estimated from samples.

The next step: the probability of a trajectory is

pθ(τ)=ρ(s0)∏tP(st+1∣st,at) πθ(at∣st)p_\theta(\tau) = \rho(s_0)\prod_t P(s_{t+1} \mid s_t, a_t)\, \pi_\theta(a_t \mid s_t)

Take the logarithm and it becomes a sum, and everything that does not depend on θ\theta disappears on differentiation — including the environment's dynamics PP. That is the whole point: we need no model of the world.

∇θJ=E[∑t∇θlog⁡πθ(at∣st) R(τ)]\nabla_\theta J = \mathbb{E}\left[\sum_t \nabla_\theta \log \pi_\theta(a_t \mid s_t)\, R(\tau)\right]

Two improvements that do not change the expectation but greatly reduce the variance:

  1. Causality. An action cannot affect what has already happened. Swap R(τ)R(\tau) for the return forwards from step tt: Gt=∑k≥tγk−trkG_t = \sum_{k \geq t}\gamma^{k-t} r_k.
  2. A baseline. For any function b(s)b(s) we have E[∇θlog⁡πθ(a∣s) b(s)]=0\mathbb{E}[\nabla_\theta \log \pi_\theta(a\mid s)\, b(s)] = 0, since ∑a∇θπθ(a∣s)=∇θ1=0\sum_a \nabla_\theta \pi_\theta(a \mid s) = \nabla_\theta 1 = 0. So we can subtract b(s)b(s) for free.

With b(s)=V(s)b(s) = V(s), Gt−V(st)G_t - V(s_t) becomes the advantage AtA_t — and then we have arrived at actor–critic.

∇θJ=E[∑t∇θlog⁡πθ(at∣st) At]\boxed{\nabla_\theta J = \mathbb{E}\left[\sum_t \nabla_\theta \log \pi_\theta(a_t \mid s_t)\, A_t\right]}

Why the variance is the problem: the estimate is built on the random outcomes of whole episodes. Without a baseline the gradient can point in different directions between two runs with the same policy. The baseline removes the part of the signal that is common to every action in a state — and that part is nothing but noise.

Code

import torch, torch.nn as nn

class Policy(nn.Module):
    def __init__(self, obs, actions, hidden=128):
        super().__init__()
        self.f = nn.Sequential(nn.Linear(obs, hidden), nn.Tanh(), nn.Linear(hidden, actions))

    def forward(self, s):
        return torch.distributions.Categorical(logits=self.f(s))

def returns(rewards, gamma=0.99):
    out, G = [], 0.0
    for r in reversed(rewards):
        G = r + gamma * G
        out.append(G)
    return list(reversed(out))

def reinforce(env, policy, opt, episodes=1000, gamma=0.99):
    for _ in range(episodes):
        s, done = env.reset()[0], False
        logp, rews = [], []
        while not done:
            d = policy(torch.as_tensor(s, dtype=torch.float32))
            a = d.sample()
            logp.append(d.log_prob(a))
            s, r, term, trunc, _ = env.step(int(a))
            rews.append(r); done = term or trunc

        G = torch.tensor(returns(rews, gamma))
        G = (G - G.mean()) / (G.std() + 1e-8)          # a baseline: greatly reduces the variance
        loss = -(torch.stack(logp) * G).sum()          # the minus sign: we are maximising J
        opt.zero_grad(); loss.backward()
        nn.utils.clip_grad_norm_(policy.parameters(), 1.0)
        opt.step()

Three things that usually go wrong:

SymptomCause
The policy quickly becomes deterministic and stops improvingthe entropy has collapsed — add an entropy bonus
No learning at allthe minus sign is missing, or the normalisation removed all the signal
Very shaky learningnormalising over a single short episode; batch several

The normalisation on line 24 is standard but not innocent: do it over a single episode and you are only comparing actions within that episode, and the information that the whole episode was bad is lost.

Mastery means

  • Derives the policy gradient theorem with the log-derivative trick
  • Implements REINFORCE with a baseline
  • Explains the variance problem and the countermeasures

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences