Reinforcement learning: learning from feedback
Understand how algorithms learn through reward and identify risks associated with incorrect reward structures.
Prerequisites
- AControlling a process with commandsrequired
Everyday explanation
How do you optimise a process? Not by dictating every step, but by rewarding outcomes. When a system happens to produce a desired result, it is registered. After many iterations, the frequency of that behaviour increases.
An AI algorithm works the same way: it performs an action, receives a score (reward) if the result was good, and adjusts its strategy to maximise the score. No one has coded how it should think — only what constitutes a good result.
This is called reinforcement learning.
Interactive
Example: Route optimisation. An algorithm must navigate from point A to point B. It can choose to go left or right.
[Start] [ ] [ ] [Goal]
Rule: +1 point when the goal is reached, 0 otherwise. Initially, it chooses randomly. Sometimes it reaches the goal → +1. In the next iteration, «right» becomes statistically more likely. After twenty runs, it chooses the direct path.
Risk: If you reward the wrong thing, the system learns the wrong thing. If you give +1 for every step towards the goal, the algorithm may learn to hover near the goal instead of completing the task — this yields a higher total score. The reward must exactly reflect the final outcome you want to achieve.
Mastery means
- Explains that reward signals determine which behaviours are reinforced
- Can describe how an agent is optimised through iteration and feedback
Sign in to do the exercises and build your mastery up.
Sources
- CS Unplugged (CC BY-SA 4.0) — CC BY-SA 4.0
- Sutton & Barto — Reinforcement Learning (free PDF), ch. 1 — free to read