Terminated vs Truncated

Learn why Gymnasium separates the done flag into terminated and truncated, and why this distinction matters for correct RL training.

By Kevin McAleer,    5 Minutes

Page last updated June 14, 2026


Cover


The Old Way: One Flag

In the previous course, the BurgerBot simulator returned a single done flag:

# The hand-built step() from the RL course
state, reward, done = world.step(action)

done was True in two situations:

  • The robot hit an obstacle (bad ending β€” episode is truly over)
  • The episode hit MAX_STEPS = 150 (time limit β€” we just stopped the episode)

Those two situations are logically different, and collapsing them into one flag causes a subtle but real training problem. Let’s look at why.

The Bootstrap Problem

Q-learning updates a value estimate based on what the agent expects to earn from the next state onward:

# The update from the RL course
# td_target = reward + gamma * max(q_table[next_state])
# q_table[state][action] += alpha * (td_target - q_table[state][action])

The term gamma * max(q_table[next_state]) is called the bootstrap: we are using our current estimate of the next state’s value to update our estimate of the current state’s value.

When the episode genuinely ends β€” the pole falls, the robot crashes, the goal is reached β€” there is no next state. The bootstrap term should be zero. That is correct.

When the episode ends only because of a time limit, the agent is actually in a perfectly valid state. The environment could continue from there. The bootstrap term should NOT be zero β€” it should reflect how much future reward the agent could still earn.

If you treat a time-limit ending the same as a genuine terminal state, you are incorrectly telling the agent β€œthere is no future value from this state” when there actually is. For small, fast-converging problems this error is often small enough to not matter. For harder problems β€” or when you are using the time limit to learn about regions of the state space the agent hasn’t reached yet β€” it introduces a real bias.

The New Way: Two Flags

Gymnasium 1.x separates the concept into two flags:

obs, reward, terminated, truncated, info = env.step(action)
  • terminated is True when the episode ends for a reason defined by the environment: the goal was reached, the agent died, the pole fell past the threshold, the robot hit a wall. This is a genuine terminal state. Bootstrap should be zero.
  • truncated is True when the episode ends because of an external limit: a time limit, a step budget, or any stopping condition imposed from outside the environment’s own rules. This is NOT a terminal state. Bootstrap should continue.

The episode is over (from the training loop’s perspective) if either is True. But a well-written RL algorithm can handle them differently.

Here is what the distinction looks like in practice:

import gymnasium as gym

env = gym.make("CartPole-v1")
obs, info = env.reset(seed=0)

for step in range(1000):
    action = env.action_space.sample()
    obs, reward, terminated, truncated, info = env.step(action)

    if terminated:
        # Pole fell β€” genuine terminal state.
        # Bootstrap = 0. No future reward possible.
        print(f"Step {step}: TERMINATED (pole fell). Resetting.")
        obs, info = env.reset()

    elif truncated:
        # Time limit hit β€” agent is still in a valid state.
        # Bootstrap could continue. Policy was cut short, not beaten.
        print(f"Step {step}: TRUNCATED (time limit). Resetting.")
        obs, info = env.reset()

env.close()

What This Means for BurgerBot

Look back at the BurgerBot simulator from the previous course. Its done flag was set True in three cases:

  1. The robot hit an obstacle (hit = True) β€” this is terminated
  2. The robot reached the goal β€” this is also terminated
  3. steps >= MAX_STEPS β€” this is truncated

The validated BurgerBotEnv you will build in lesson 8 handles all three correctly:

# From BurgerBotEnv.step()
reached = (self.row, self.col) == GOAL

if hit:
    reward = -10.0
elif reached:
    reward = 50.0
elif moved:
    reward = 1.0
else:
    reward = -0.1

terminated = reached or hit          # genuine endings
truncated = self.steps >= self.max_steps   # time limit only

return self._obs(), reward, terminated, truncated, {}

A hit gives terminated=True, truncated=False. A time limit gives terminated=False, truncated=True. Both end the episode from the training loop’s point of view, but Stable-Baselines3 (and any correct RL implementation) handles them differently.

The Simple Training Loop Pattern

For most hand-written training loops, you handle both flags the same way at the episode level:

done = terminated or truncated
if done:
    obs, info = env.reset()

This is fine for tabular Q-learning where the distinction has little practical effect. When you move to SB3 in lessons 12 and 13, the library handles the distinction for you internally β€” which is one more reason to use a standard API.

Try It Yourself

  1. Run the code snippet above with render_mode omitted. After 200 steps, count how many TERMINATED events and how many TRUNCATED events you saw. CartPole’s default time limit is 500 steps per episode β€” does this match your truncation count?
  2. Change the random policy to action = 0 (always push left). Does the pole fall faster, making TERMINATED more common?
  3. Read the CartPole source code to find where terminated and truncated are set. You can find it by running print(env.spec.entry_point) β€” it will tell you the module path.

Common Issues

Problem: Code written for old OpenAI Gym gives ValueError: too many values to unpack Solution: Change obs, reward, done, info = env.step(action) to obs, reward, terminated, truncated, info = env.step(action). Why: Gymnasium 1.x returns five values from step(). The old Gym returned four.

Problem: AttributeError: 'tuple' object has no attribute 'dtype' Solution: You are probably unpacking reset() as obs = env.reset() instead of obs, info = env.reset(). Why: Gymnasium 1.x reset() returns a two-tuple. Assigning it to a single variable gives you the tuple, not the observation.

< Previous Next >

You can use the arrows  ← β†’ on your keyboard to navigate between lessons.


What are you looking for?
Watch Videos Get Ideas Learn Something Read a Review Read the Blog Search
... Z Z z