The goal E University
Train an agent with reward
From bandits to policy gradient: build an environment, let an agent learn from reward, and understand why RL is hard.
- Knowledge nodes
- 31
- From zero
- about 18 h
- Labs
- 6
See what you already know — no account
The diagnostic removes what you already know, so your path is usually much shorter.
What you can do afterwards
Labs along the way
You write the code. Tests you cannot see decide whether it holds up.
Lab: gradient descent from scratchDin the browser · about 50 minLab: lists, loops and dictionariesCin the browser · about 40 minLab: mean, median and standard deviation by handCin the browser · about 40 minLab: simulate dice and compare with the theoryCin the browser · about 40 minLab: the best line with least squaresCin the browser · about 40 minLab: your first Python functionsCin the browser · about 30 min
The whole path
Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map