The grid world: your first RL agent
Be able to implement an agent that learns to find its way to the goal in a grid.
Prerequisites
- CTry something new or play it safe?required
- CPython — lists, loops and dictionariesrequired
Intuition
The grid world is RL's equivalent of «hello world». A small grid, an agent, a goal.
. . . G G = the goal (+1)
. # . X X = the trap (−1)
A . . . # = a wall
All four concepts are here, in their simplest form:
| Concept | In the grid |
|---|---|
| State | which square the agent is standing in |
| Action | up, down, left, right |
| Reward | +1 at the goal, −1 in the trap, otherwise a small minus per step |
| Episode | from the start until the agent reaches the goal or the trap |
The agent knows nothing at the start. It feels its way, gets points, and slowly remembers which squares are worth heading for.
The small minus per step matters more than it sounds: without it the length of the route makes no difference, and the agent can learn to wander around forever before going to the goal.
Code
import random
WIDTH, HEIGHT = 4, 3
WALLS = {(1, 1)}
GOAL, TRAP = (0, 3), (1, 3)
ACTIONS = ["up", "down", "left", "right"]
DELTA = {"up": (-1, 0), "down": (1, 0), "left": (0, -1), "right": (0, 1)}
def move(pos, action):
dr, dc = DELTA[action]
new = (pos[0] + dr, pos[1] + dc)
if not (0 <= new[0] < HEIGHT and 0 <= new[1] < WIDTH) or new in WALLS:
return pos # walk into a wall → stay put
return new
def reward(pos):
if pos == GOAL:
return 1.0, True
if pos == TRAP:
return -1.0, True
return -0.04, False # a small minus per step → short routes pay off
Q = {(r, c, a): 0.0 for r in range(HEIGHT) for c in range(WIDTH) for a in ACTIONS}
def best(pos):
return max(ACTIONS, key=lambda a: Q[(pos[0], pos[1], a)])
rng = random.Random(0)
for episode in range(2000):
eps = max(0.05, 1.0 - episode / 1000) # explore a lot first, then less
pos, steps = (2, 0), 0
while steps < 100:
a = rng.choice(ACTIONS) if rng.random() < eps else best(pos)
new = move(pos, a)
r, done = reward(new)
best_next = 0.0 if done else max(Q[(new[0], new[1], x)] for x in ACTIONS)
Q[(pos[0], pos[1], a)] += 0.1 * (r + 0.95 * best_next - Q[(pos[0], pos[1], a)])
pos, steps = new, steps + 1
if done:
break
LABEL = {GOAL: "GOAL ", TRAP: "TRAP ", **{w: " # " for w in WALLS}}
for r in range(HEIGHT):
print(" ".join(LABEL.get((r, c), f"{best((r, c)):^5.5}") for c in range(WIDTH)))
# right right right GOAL
# up # up TRAP
# up right up left
The arrows point towards the goal and away from the trap — without anyone having said where they are. The agent has only tried things, got points, and remembered.
Interactive
Four experiments that change what the agent learns. Run the base version first, look at the arrows, then change one thing at a time.
| Change | What happens |
|---|---|
Remove the per-step penalty (-0.04 → 0.0) | the arrows stop pointing the direct way; every route is equally good |
Raise the penalty to -0.5 | the agent walks straight into the trap — dying quickly beats walking far |
Set eps = 0.0 | the agent gets stuck on the first route it found, often a bad one |
Set 0.95 → 0.5 | the agent becomes short-sighted and does not see goals that are far off |
The second row is the most instructive. The reward is not a wish — it is a definition. Punish every step too hard and you have actually asked the agent to finish quickly, and the trap finishes quickly.
It is the same phenomenon that makes reward design hard in real systems, and it is clearest here in the grid where you can see the whole solution at once.
Mastery means
- Describes the state, the action and the reward in a grid
- Implements movement with walls
- Explains how the reward steers what the agent learns
Sign in to do the exercises and build your mastery up.
Sources
- Sutton & Barto — Reinforcement Learning: An Introduction (2:a uppl.) — free to read online (authors' edition)
- Gymnasium — dokumentation (MIT) — MIT
- The Python documentation (PSF licence) — PSF