Skip to content
AI-grafen
EUniversityReinforcement learning· about 60 min· fundamentals that rarely change· verified 2026-09-20· EN

Reinforcement learning — the basics

Be able to explain the agent, the environment, the reward and the policy, and implement a simple policy gradient.

Prerequisites

Intuition

In supervised learning the model gets the answer for every example. In RL the agent only gets a reward — often late, often sparse — and has to work out for itself which actions led there.

The vocabulary: the state s (what the agent sees), the action a, the reward r, the policy π(a|s) (the agent's behaviour, often a neural network), the return G = the sum of the (discounted) rewards from now on.

Policy gradient (REINFORCE): run an episode, and for every action: make it more likely if the return after it was high, less likely if it was low. The gradient is ∑ₜ ∇log π(aₜ|sₜ)·Gₜ. Simple, but noisy — one episode says little about which action was good (credit assignment). You subtract a baseline (the average return) to reduce the variance.

RLHF for language models is exactly this: the state = the prompt so far, the action = the next token, the reward = a reward model's score on the whole answer.

Code

import numpy as np
rng = np.random.default_rng(0)

# The environment: go from 0 to 5 on a line; action 0 = left, 1 = right; reward 1 at the goal, 20 steps max
def episode(theta):
    s, traj = 0, []
    for _ in range(20):
        p = 1 / (1 + np.exp(-theta[s]))           # π(right | s)
        a = int(rng.random() < p)
        traj.append((s, a, p))
        s = min(5, max(0, s + (1 if a else -1)))
        if s == 5: return traj, 1.0
    return traj, 0.0

theta = np.zeros(6); lr = 0.5; baseline = 0.0
for it in range(300):
    traj, G = episode(theta)
    adv = G - baseline; baseline = 0.9 * baseline + 0.1 * G
    for s, a, p in traj:
        theta[s] += lr * adv * ((a - p))            # ∇log π for a Bernoulli/sigmoid
print(np.round(1 / (1 + np.exp(-theta)), 2))        # → close to 1 everywhere: «go right»

Formal

The objective: maximise J(θ)=Eτ∼πθ[G(τ)]J(\theta) = \mathbb E_{\tau\sim\pi_\theta}[G(\tau)]. The policy gradient theorem: ∇θJ=E[∑t∇θlog⁡πθ(at∣st) (Gt−b(st))]\nabla_\theta J = \mathbb E\big[\sum_t \nabla_\theta\log\pi_\theta(a_t\mid s_t)\,(G_t - b(s_t))\big], where the baseline bb does not change the expectation but lowers the variance. With b=Vπ(s)b = V^\pi(s), Gt−bG_t - b becomes an estimate of the advantage A(s,a)A(s,a) — the basis of actor–critic and PPO. Discounting Gt=∑kγkrt+kG_t = \sum_k \gamma^k r_{t+k} with γ∈[0,1)\gamma\in[0,1) gives a finite return and a shorter credit-assignment horizon.

Mastery means

  • Defines the agent, the environment, the state, the action, the reward and the policy
  • Implements REINFORCE on a simple environment
  • Explains why RL is hard: credit assignment, variance, exploration

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences