Skip to content
AI-grafen

The goal E University

Train an agent with reward

From bandits to policy gradient: build an environment, let an agent learn from reward, and understand why RL is hard.

Knowledge nodes
31
From zero
about 18 h
Labs
6
See what you already know — no account

The diagnostic removes what you already know, so your path is usually much shorter.

What you can do afterwards

Labs along the way

You write the code. Tests you cannot see decide whether it holds up.

The whole path

Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map