KevsRobots Learning Platform
24% Percent Complete
By Kevin McAleer, 5 Minutes
Page last updated June 14, 2026

In the previous course, the BurgerBot simulator returned a single done flag:
# The hand-built step() from the RL course
state, reward, done = world.step(action)
done was True in two situations:
MAX_STEPS = 150 (time limit β we just stopped the episode)Those two situations are logically different, and collapsing them into one flag causes a subtle but real training problem. Letβs look at why.
Q-learning updates a value estimate based on what the agent expects to earn from the next state onward:
# The update from the RL course
# td_target = reward + gamma * max(q_table[next_state])
# q_table[state][action] += alpha * (td_target - q_table[state][action])
The term gamma * max(q_table[next_state]) is called the bootstrap: we are using our current estimate of the next stateβs value to update our estimate of the current stateβs value.
When the episode genuinely ends β the pole falls, the robot crashes, the goal is reached β there is no next state. The bootstrap term should be zero. That is correct.
When the episode ends only because of a time limit, the agent is actually in a perfectly valid state. The environment could continue from there. The bootstrap term should NOT be zero β it should reflect how much future reward the agent could still earn.
If you treat a time-limit ending the same as a genuine terminal state, you are incorrectly telling the agent βthere is no future value from this stateβ when there actually is. For small, fast-converging problems this error is often small enough to not matter. For harder problems β or when you are using the time limit to learn about regions of the state space the agent hasnβt reached yet β it introduces a real bias.
Gymnasium 1.x separates the concept into two flags:
obs, reward, terminated, truncated, info = env.step(action)
terminated is True when the episode ends for a reason defined by the environment: the goal was reached, the agent died, the pole fell past the threshold, the robot hit a wall. This is a genuine terminal state. Bootstrap should be zero.truncated is True when the episode ends because of an external limit: a time limit, a step budget, or any stopping condition imposed from outside the environmentβs own rules. This is NOT a terminal state. Bootstrap should continue.The episode is over (from the training loopβs perspective) if either is True. But a well-written RL algorithm can handle them differently.
Here is what the distinction looks like in practice:
import gymnasium as gym
env = gym.make("CartPole-v1")
obs, info = env.reset(seed=0)
for step in range(1000):
action = env.action_space.sample()
obs, reward, terminated, truncated, info = env.step(action)
if terminated:
# Pole fell β genuine terminal state.
# Bootstrap = 0. No future reward possible.
print(f"Step {step}: TERMINATED (pole fell). Resetting.")
obs, info = env.reset()
elif truncated:
# Time limit hit β agent is still in a valid state.
# Bootstrap could continue. Policy was cut short, not beaten.
print(f"Step {step}: TRUNCATED (time limit). Resetting.")
obs, info = env.reset()
env.close()
Look back at the BurgerBot simulator from the previous course. Its done flag was set True in three cases:
terminatedterminatedsteps >= MAX_STEPS β this is truncatedThe validated BurgerBotEnv you will build in lesson 8 handles all three correctly:
# From BurgerBotEnv.step()
reached = (self.row, self.col) == GOAL
if hit:
reward = -10.0
elif reached:
reward = 50.0
elif moved:
reward = 1.0
else:
reward = -0.1
terminated = reached or hit # genuine endings
truncated = self.steps >= self.max_steps # time limit only
return self._obs(), reward, terminated, truncated, {}
A hit gives terminated=True, truncated=False. A time limit gives terminated=False, truncated=True. Both end the episode from the training loopβs point of view, but Stable-Baselines3 (and any correct RL implementation) handles them differently.
For most hand-written training loops, you handle both flags the same way at the episode level:
done = terminated or truncated
if done:
obs, info = env.reset()
This is fine for tabular Q-learning where the distinction has little practical effect. When you move to SB3 in lessons 12 and 13, the library handles the distinction for you internally β which is one more reason to use a standard API.
render_mode omitted. After 200 steps, count how many TERMINATED events and how many TRUNCATED events you saw. CartPoleβs default time limit is 500 steps per episode β does this match your truncation count?action = 0 (always push left). Does the pole fall faster, making TERMINATED more common?terminated and truncated are set. You can find it by running print(env.spec.entry_point) β it will tell you the module path.Problem: Code written for old OpenAI Gym gives ValueError: too many values to unpack
Solution: Change obs, reward, done, info = env.step(action) to obs, reward, terminated, truncated, info = env.step(action).
Why: Gymnasium 1.x returns five values from step(). The old Gym returned four.
Problem: AttributeError: 'tuple' object has no attribute 'dtype'
Solution: You are probably unpacking reset() as obs = env.reset() instead of obs, info = env.reset().
Why: Gymnasium 1.x reset() returns a two-tuple. Assigning it to a single variable gives you the tuple, not the observation.
You can use the arrows β β on your keyboard to navigate between lessons.
Comments