Reinforcement learning — the basics
Be able to explain the agent, the environment, the reward and the policy, and implement a simple policy gradient.
Prerequisites
- CProbability — the basicsrequiredPractise in Mattegrafen ↗
- DGradient descentrequired
Intuition
In supervised learning the model gets the answer for every example. In RL the agent only gets a reward — often late, often sparse — and has to work out for itself which actions led there.
The vocabulary: the state s (what the agent sees), the action a, the reward r, the policy π(a|s) (the agent's behaviour, often a neural network), the return G = the sum of the (discounted) rewards from now on.
Policy gradient (REINFORCE): run an episode, and for every action: make it more likely if the return after it was high, less likely if it was low. The gradient is ∑ₜ ∇log π(aₜ|sₜ)·Gₜ. Simple, but noisy — one episode says little about which action was good (credit assignment). You subtract a baseline (the average return) to reduce the variance.
RLHF for language models is exactly this: the state = the prompt so far, the action = the next token, the reward = a reward model's score on the whole answer.
Code
import numpy as np
rng = np.random.default_rng(0)
# The environment: go from 0 to 5 on a line; action 0 = left, 1 = right; reward 1 at the goal, 20 steps max
def episode(theta):
s, traj = 0, []
for _ in range(20):
p = 1 / (1 + np.exp(-theta[s])) # π(right | s)
a = int(rng.random() < p)
traj.append((s, a, p))
s = min(5, max(0, s + (1 if a else -1)))
if s == 5: return traj, 1.0
return traj, 0.0
theta = np.zeros(6); lr = 0.5; baseline = 0.0
for it in range(300):
traj, G = episode(theta)
adv = G - baseline; baseline = 0.9 * baseline + 0.1 * G
for s, a, p in traj:
theta[s] += lr * adv * ((a - p)) # ∇log π for a Bernoulli/sigmoid
print(np.round(1 / (1 + np.exp(-theta)), 2)) # → close to 1 everywhere: «go right»
Formal
The objective: maximise . The policy gradient theorem: , where the baseline does not change the expectation but lowers the variance. With , becomes an estimate of the advantage — the basis of actor–critic and PPO. Discounting with gives a finite return and a shorter credit-assignment horizon.
Mastery means
- Defines the agent, the environment, the state, the action, the reward and the policy
- Implements REINFORCE on a simple environment
- Explains why RL is hard: credit assignment, variance, exploration
Sign in to do the exercises and build your mastery up.
Sources
- Sutton & Barto — Reinforcement Learning: An Introduction (free PDF) — free to read
- arXiv — Training language models to follow instructions with human feedback — arXiv (open access; licence per article)