EUniversityReinforcement learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN
Environments and simulation
Be able to build a simple environment with a Gym-like interface.
Prerequisites
- DPython — classes and objectsrequired
Intuition
An RL environment is an object with two methods: reset() → the first observation; step(a) → (observation, reward, done, info). That is the Gymnasium interface, and everything — from CartPole to a trading simulator or a «pupil answering exercises» — can be expressed that way.
The design choices that decide whether the agent learns anything:
- The observation: enough to make a decision (Markov), no more.
- The reward: dense enough to give a signal, but what you actually want. Reward «being near the goal» and the agent learns to circle near the goal.
- Termination: the goal reached, a time limit, a failure.
- Determinism:
reset(seed=…)for reproducible evaluation.
Test the environment before the agent: a random agent should give plausible rewards; a hand-written «correct» behaviour should give a high reward. Many RL bugs are environment bugs.
Code
import numpy as np
class Grid:
"""A 4×4 grid: start (0,0), goal (3,3), a hole at (1,1). Actions 0–3 = up/right/down/left."""
def __init__(self, max_steps=30):
self.max_steps = max_steps; self.rng = np.random.default_rng()
def reset(self, seed=None):
self.rng = np.random.default_rng(seed); self.pos = np.array([0, 0]); self.t = 0
return self._obs()
def _obs(self):
return int(self.pos[0] * 4 + self.pos[1])
def step(self, a):
d = [(-1, 0), (0, 1), (1, 0), (0, -1)][a]
self.pos = np.clip(self.pos + d, 0, 3); self.t += 1
if (self.pos == [1, 1]).all(): return self._obs(), -1.0, True, {"why": "hole"}
if (self.pos == [3, 3]).all(): return self._obs(), 1.0, True, {"why": "goal"}
return self._obs(), -0.01, self.t >= self.max_steps, {}
env = Grid(); obs = env.reset(seed=0); total = 0
for _ in range(30):
obs, r, done, info = env.step(env.rng.integers(4)); total += r
if done: break
print(total, info)
Small negative step rewards (−0.01) encourage short routes without dominating the goal reward.
Mastery means
- Builds an environment with a reset/step interface
- Defines the observation space, the action space and the reward
- Tests the environment with a random agent and a deterministic seed
Sign in to do the exercises and build your mastery up.