Skip to content
AI-grafen
EUniversityReinforcement learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Environments and simulation

Be able to build a simple environment with a Gym-like interface.

Prerequisites

Intuition

An RL environment is an object with two methods: reset() → the first observation; step(a) → (observation, reward, done, info). That is the Gymnasium interface, and everything — from CartPole to a trading simulator or a «pupil answering exercises» — can be expressed that way.

The design choices that decide whether the agent learns anything:

  • The observation: enough to make a decision (Markov), no more.
  • The reward: dense enough to give a signal, but what you actually want. Reward «being near the goal» and the agent learns to circle near the goal.
  • Termination: the goal reached, a time limit, a failure.
  • Determinism: reset(seed=…) for reproducible evaluation.

Test the environment before the agent: a random agent should give plausible rewards; a hand-written «correct» behaviour should give a high reward. Many RL bugs are environment bugs.

Code

import numpy as np

class Grid:
    """A 4×4 grid: start (0,0), goal (3,3), a hole at (1,1). Actions 0–3 = up/right/down/left."""
    def __init__(self, max_steps=30):
        self.max_steps = max_steps; self.rng = np.random.default_rng()
    def reset(self, seed=None):
        self.rng = np.random.default_rng(seed); self.pos = np.array([0, 0]); self.t = 0
        return self._obs()
    def _obs(self):
        return int(self.pos[0] * 4 + self.pos[1])
    def step(self, a):
        d = [(-1, 0), (0, 1), (1, 0), (0, -1)][a]
        self.pos = np.clip(self.pos + d, 0, 3); self.t += 1
        if (self.pos == [1, 1]).all(): return self._obs(), -1.0, True, {"why": "hole"}
        if (self.pos == [3, 3]).all(): return self._obs(), 1.0, True, {"why": "goal"}
        return self._obs(), -0.01, self.t >= self.max_steps, {}

env = Grid(); obs = env.reset(seed=0); total = 0
for _ in range(30):
    obs, r, done, info = env.step(env.rng.integers(4)); total += r
    if done: break
print(total, info)

Small negative step rewards (−0.01) encourage short routes without dominating the goal reward.

Mastery means

  • Builds an environment with a reset/step interface
  • Defines the observation space, the action space and the reward
  • Tests the environment with a random agent and a deterministic seed

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences