KevsRobots Learning Platform
30% Percent Complete
By Kevin McAleer, 5 Minutes
Page last updated June 14, 2026

When you wrote the BurgerBot Q-learning agent by hand, your choose_action function knew the exact structure of the state: a tuple of (row, col, heading, sensor). It knew that row ranged from 0 to 6, that heading was one of four string values, and that sensor had three levels. You hard-coded that knowledge.
A generic algorithm β one that should work on CartPole, BurgerBot, and a simulated robot arm without modification β cannot have that knowledge baked in. It needs a way to ask: βWhat does the observation look like? What actions are available?β
That is what observation_space and action_space are for. They are machine-readable descriptions of the shape and range of the data the environment produces and accepts.
Gymnasium defines many space types. Three cover the vast majority of environments you will encounter.
A single integer chosen from {0, 1, 2, ..., n-1}.
from gymnasium import spaces
# Four possible actions: 0, 1, 2, 3
action_space = spaces.Discrete(4)
print(action_space.n) # 4
print(action_space.sample()) # random integer in {0, 1, 2, 3}
print(action_space.contains(2)) # True
print(action_space.contains(4)) # False β 4 is out of range
BurgerBot uses Discrete(4) for actions: 0=forward, 1=turn_left, 2=turn_right, 3=stop.
A continuous n-dimensional array, where each dimension has its own lower and upper bound.
import numpy as np
from gymnasium import spaces
# CartPole's observation: 4 continuous values
# [cart_position, cart_velocity, pole_angle, pole_angular_velocity]
obs_space = spaces.Box(
low=np.array([-4.8, -np.inf, -0.418, -np.inf]),
high=np.array([4.8, np.inf, 0.418, np.inf]),
dtype=np.float32
)
print(obs_space.shape) # (4,)
print(obs_space.sample()) # random float array β will be in range
print(obs_space.contains(np.zeros(4, dtype=np.float32))) # True
Box is what you get with continuous sensor data, camera images (a Box of shape (H, W, 3)), joint angles, velocities, and most real-world robotics observations.
An array of independent integers, each from its own range. This is what BurgerBot uses for observations.
from gymnasium import spaces
# BurgerBot observation: [row, col, heading_index, sensor_reading]
# row: 0-6, col: 0-6, heading: 0-3, sensor: 0-2
obs_space = spaces.MultiDiscrete([7, 7, 4, 3])
print(obs_space.nvec) # array([7, 7, 4, 3])
print(obs_space.sample()) # e.g. array([3, 5, 1, 2])
print(obs_space.contains(np.array([6, 6, 3, 2]))) # True
print(obs_space.contains(np.array([7, 0, 0, 0]))) # False β row 7 out of range
The total number of possible observations is 7 * 7 * 4 * 3 = 588. This matches exactly what the previous course calculated when discussing the state-space size.
Here is a simplified version of what Stable-Baselines3 does internally when you call model = DQN("MlpPolicy", env):
# The algorithm inspects the spaces without knowing what the environment is.
n_actions = env.action_space.n # 4 for BurgerBot
obs_shape = env.observation_space.shape # (4,) for CartPole
# It builds a neural network whose input size matches the observation shape
# and whose output size matches the number of actions.
# It never reads the actual environment code.
This is the power of the contract. The algorithm does not know if it is training on CartPole or BurgerBot. It only knows the shape of the data coming in and the range of actions going out.
You can always inspect any Gymnasium environmentβs spaces:
import gymnasium as gym
env = gym.make("CartPole-v1")
print("Observation space:", env.observation_space)
print("Action space: ", env.action_space)
print("Obs shape: ", env.observation_space.shape)
print("Obs dtype: ", env.observation_space.dtype)
print("Num actions: ", env.action_space.n)
env.close()
Output:
Observation space: Box([-4.8 -inf -0.4188 -inf], [4.8 inf 0.4188 inf], (4,), float32)
Action space: Discrete(2)
Obs shape: (4,)
Obs dtype: float32
Num actions: 2
space.sample() is used in two places in RL code:
action = env.action_space.sample() gives you a legal random action without knowing the environmentβs internals.env.action_space.sample() instead of querying the Q-table.import random
epsilon = 0.3
if random.random() < epsilon:
# Explore: pick a random legal action
action = env.action_space.sample()
else:
# Exploit: pick the best known action
action = best_action_from_q_table(state)
This replaces the random.choice(ACTIONS) call from the previous course with something that works on any environment, regardless of how many actions it has.
Box space for a 64x64 greyscale image observation: values between 0 and 255, dtype=np.uint8. What is space.shape? How many possible observations exist? (You donβt need to compute the exact number β just express it symbolically.)MultiDiscrete([3, 3]) space representing a 3x3 grid position. How many distinct positions can the agent be in?LunarLander-v3 environmentβs spaces (use gym.make("LunarLander-v3")). Is the action space discrete or continuous? How many observations does the agent get?Problem: AssertionError: observation is not in observation_space from check_env
Solution: Your reset() or step() returned an observation with the wrong dtype or out-of-range values.
Why: NumPy arrays have strict dtypes. np.array([1, 2, 3]) defaults to int64, but MultiDiscrete expects int64 too β usually fine. Box with dtype=float32 rejects a float64 array. Use dtype=np.int64 or dtype=np.float32 explicitly in your environment.
Problem: AttributeError: 'Box' object has no attribute 'n'
Solution: .n is a Discrete attribute. Use env.action_space.shape or env.action_space.low for Box spaces.
Why: Different space types have different attributes. Always check the space type before accessing type-specific properties.
You can use the arrows β β on your keyboard to navigate between lessons.
Comments