EUniversityReinforcement learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN
Multi-armed bandits
Be able to implement epsilon-greedy and Thompson sampling — the method that chooses the explanation style in the platform.
Prerequisites
Intuition
You have k slot machines («arms») with unknown win rates. Every pull: pull an arm, see the reward. The goal: maximise the total reward. The problem: to know which one is best you have to explore bad arms; to earn you have to exploit the best one.
- ε-greedy: with probability ε pull a random arm, otherwise the one with the highest estimated mean. Simple; explores just as much forever.
- Thompson sampling: keep a probability distribution (Beta) over each arm's rate; draw one sample per arm and pick the highest. It automatically explores more where the uncertainty is large. Close to optimal in practice.
- UCB: pick the arm with the highest mean plus a bonus for uncertainty.
Regret = what you lost compared with always having pulled the best arm. Good algorithms have regret that grows like log(t).
AI-grafen uses this to choose which type of explanation works best for a user: every explanation depth is an arm, and «the user solved the exercise» is the reward.
Code
import numpy as np
rng = np.random.default_rng(0)
p_true = np.array([0.30, 0.45, 0.60, 0.40]) # unknown to the agent
def run(choose, T=5000):
n = np.zeros(4); s = np.zeros(4); regret = 0.0
for t in range(T):
a = choose(n, s, t)
r = rng.random() < p_true[a]
n[a] += 1; s[a] += r; regret += p_true.max() - p_true[a]
return regret
def eps_greedy(eps=0.1):
def f(n, s, t):
if rng.random() < eps or n.min() == 0: return rng.integers(4)
return int(np.argmax(s / n))
return f
def thompson(n, s, t):
return int(np.argmax(rng.beta(s + 1, n - s + 1)))
print("eps-greedy", round(run(eps_greedy()), 1)) # ≈ 90
print("thompson ", round(run(thompson), 1)) # ≈ 25
Mastery means
- Implements epsilon-greedy and Thompson sampling
- Explains the explore/exploit trade-off and regret
- Sees where bandits are used in products (and in the platform)
Sign in to do the exercises and build your mastery up.
Sources
- Sutton & Barto — Reinforcement Learning: An Introduction (free PDF), ch. 2 — free to read
- arXiv — A Tutorial on Thompson Sampling — arXiv (open access; licence per article)