Skip to content
AI-grafen
DAI developerReinforcement learning· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

The grid world: your first RL agent

Be able to implement an agent that learns to find its way to the goal in a grid.

Prerequisites

Intuition

The grid world is RL's equivalent of «hello world». A small grid, an agent, a goal.

. . . G      G = the goal (+1)
. # . X      X = the trap (−1)
A . . .      # = a wall

All four concepts are here, in their simplest form:

ConceptIn the grid
Statewhich square the agent is standing in
Actionup, down, left, right
Reward+1 at the goal, −1 in the trap, otherwise a small minus per step
Episodefrom the start until the agent reaches the goal or the trap

The agent knows nothing at the start. It feels its way, gets points, and slowly remembers which squares are worth heading for.

The small minus per step matters more than it sounds: without it the length of the route makes no difference, and the agent can learn to wander around forever before going to the goal.

Code

import random

WIDTH, HEIGHT = 4, 3
WALLS = {(1, 1)}
GOAL, TRAP = (0, 3), (1, 3)
ACTIONS = ["up", "down", "left", "right"]
DELTA = {"up": (-1, 0), "down": (1, 0), "left": (0, -1), "right": (0, 1)}

def move(pos, action):
    dr, dc = DELTA[action]
    new = (pos[0] + dr, pos[1] + dc)
    if not (0 <= new[0] < HEIGHT and 0 <= new[1] < WIDTH) or new in WALLS:
        return pos                       # walk into a wall → stay put
    return new

def reward(pos):
    if pos == GOAL:
        return 1.0, True
    if pos == TRAP:
        return -1.0, True
    return -0.04, False                  # a small minus per step → short routes pay off

Q = {(r, c, a): 0.0 for r in range(HEIGHT) for c in range(WIDTH) for a in ACTIONS}

def best(pos):
    return max(ACTIONS, key=lambda a: Q[(pos[0], pos[1], a)])

rng = random.Random(0)
for episode in range(2000):
    eps = max(0.05, 1.0 - episode / 1000)         # explore a lot first, then less
    pos, steps = (2, 0), 0
    while steps < 100:
        a = rng.choice(ACTIONS) if rng.random() < eps else best(pos)
        new = move(pos, a)
        r, done = reward(new)
        best_next = 0.0 if done else max(Q[(new[0], new[1], x)] for x in ACTIONS)
        Q[(pos[0], pos[1], a)] += 0.1 * (r + 0.95 * best_next - Q[(pos[0], pos[1], a)])
        pos, steps = new, steps + 1
        if done:
            break

LABEL = {GOAL: "GOAL ", TRAP: "TRAP ", **{w: "  #  " for w in WALLS}}
for r in range(HEIGHT):
    print(" ".join(LABEL.get((r, c), f"{best((r, c)):^5.5}") for c in range(WIDTH)))
# right right right  GOAL
# up      #    up    TRAP
# up    right  up    left

The arrows point towards the goal and away from the trap — without anyone having said where they are. The agent has only tried things, got points, and remembered.

Interactive

Four experiments that change what the agent learns. Run the base version first, look at the arrows, then change one thing at a time.

ChangeWhat happens
Remove the per-step penalty (-0.04 → 0.0)the arrows stop pointing the direct way; every route is equally good
Raise the penalty to -0.5the agent walks straight into the trap — dying quickly beats walking far
Set eps = 0.0the agent gets stuck on the first route it found, often a bad one
Set 0.95 → 0.5the agent becomes short-sighted and does not see goals that are far off

The second row is the most instructive. The reward is not a wish — it is a definition. Punish every step too hard and you have actually asked the agent to finish quickly, and the trap finishes quickly.

It is the same phenomenon that makes reward design hard in real systems, and it is clearest here in the grid where you can see the whole solution at once.

Mastery means

  • Describes the state, the action and the reward in a grid
  • Implements movement with walls
  • Explains how the reward steers what the agent learns

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences