KevsRobots Learning Platform
54% Percent Complete
By Kevin McAleer, 6 Minutes
Page last updated June 13, 2026

Imagine you’re a robot with a half-learned Q-table. In your current state, turn_left looks promising (Q = 1.2) and forward looks risky (Q = -0.5). Should you always turn left?
Not necessarily. The Q-values are estimates based on experience so far — and early in training, there hasn’t been much experience. If you always pick the highest Q-value, you might be committing to a mediocre strategy before you’ve had a chance to discover better ones. But if you always pick randomly, you never benefit from what you’ve already learned.
This is the exploration-exploitation dilemma, and it sits at the heart of every reinforcement learning algorithm.
The most popular solution is beautifully simple: most of the time, pick the best-known action (exploit); occasionally, pick a random action (explore).
ε (epsilon) is the probability of choosing a random action:
import random
def choose_action(q_table, state, epsilon):
"""
ε-greedy action selection.
With probability epsilon, choose a random action (explore).
With probability 1-epsilon, choose the action with the
highest Q-value (exploit).
Args:
q_table (dict): current Q-table (built using lesson 7's defaultdict)
state (tuple): current state
epsilon (float): exploration probability, between 0 and 1
Returns:
str: the chosen action
"""
if random.random() < epsilon:
# Explore: pick a uniformly random action
return random.choice(ACTIONS)
else:
# Exploit: pick the action with the highest Q-value
action_values = q_table[state]
return max(action_values, key=lambda a: action_values[a])
With epsilon = 0.1, the agent explores 10% of the time and exploits its knowledge 90% of the time. This single parameter controls the entire balance.
At the start of training, the Q-table is all zeros — the agent knows nothing. Exploration is valuable then because it generates diverse experience. But late in training, the Q-table holds hard-won knowledge. Exploring randomly at that point just wastes time.
Epsilon decay addresses this by gradually reducing epsilon over the course of training:
def decay_epsilon(epsilon, decay_rate, epsilon_min):
"""
Reduce epsilon after each episode.
Multiply epsilon by decay_rate (e.g. 0.995), but never let
it fall below epsilon_min so the agent keeps a small amount
of curiosity even when fully trained.
Args:
epsilon (float): current exploration probability
decay_rate (float): multiplicative decay per episode (e.g. 0.995)
epsilon_min (float): lowest allowed epsilon (e.g. 0.01)
Returns:
float: new (smaller) epsilon
"""
return max(epsilon_min, epsilon * decay_rate)
# Typical training run
epsilon = 1.0 # start fully random (explore everything)
decay_rate = 0.995 # shrink epsilon by 0.5% each episode
epsilon_min = 0.01 # keep at least 1% exploration forever
for episode in range(500):
# ... run episode using choose_action(q_table, state, epsilon) ...
epsilon = decay_epsilon(epsilon, decay_rate, epsilon_min)
# After 500 episodes:
# epsilon ≈ 1.0 * (0.995 ** 500) ≈ 0.082
# Roughly 8% exploration by the end — still a little curious!
Let’s trace how epsilon evolves over time:
# Print epsilon at key checkpoints
epsilon = 1.0
decay_rate = 0.995
epsilon_min = 0.01
checkpoints = [0, 50, 100, 200, 300, 500, 750, 1000]
ep = 0
for episode in range(1001):
if episode in checkpoints:
print(f"Episode {episode:4d}: epsilon = {epsilon:.4f}")
epsilon = max(epsilon_min, epsilon * decay_rate)
# Output:
# Episode 0: epsilon = 1.0000
# Episode 50: epsilon = 0.7783
# Episode 100: epsilon = 0.6058
# Episode 200: epsilon = 0.3670
# Episode 300: epsilon = 0.2222
# Episode 500: epsilon = 0.0816
# Episode 750: epsilon = 0.0236
# Episode 1000: epsilon = 0.0100 (clamped at epsilon_min)
The agent starts out almost completely random, gradually transitions to mostly exploiting its knowledge, and settles into a 99% exploit / 1% explore regime.
Let’s trace through a few steps of the learning loop to see everything working together. We’ll use the Q-table from lesson 7:
from collections import defaultdict
import random
ACTIONS = ["forward", "turn_left", "turn_right", "stop"]
def make_q_table():
return defaultdict(lambda: {action: 0.0 for action in ACTIONS})
def choose_action(q_table, state, epsilon):
if random.random() < epsilon:
return random.choice(ACTIONS)
action_values = q_table[state]
return max(action_values, key=lambda a: action_values[a])
# Set up
q_table = make_q_table()
epsilon = 0.8 # mostly exploring early on
# Simulate a few decisions
random.seed(42)
state = (1, 1, "east", "far")
for step in range(6):
roll = random.random()
if roll < epsilon:
chosen = random.choice(ACTIONS)
mode = "EXPLORE"
else:
chosen = max(q_table[state], key=lambda a: q_table[state][a])
mode = "EXPLOIT"
print(f"Step {step}: roll={roll:.3f} ({mode}) -> {chosen}")
Notice that when all Q-values are zero, exploit and explore give the same result (any action is equally “best”). The difference becomes meaningful once some episodes have run and the Q-values have diverged.
The right decay rate depends on the task:
# For a small grid world (5x5, fast to explore):
epsilon_start = 1.0
decay_rate = 0.99 # faster decay, most exploration done in ~200 episodes
epsilon_min = 0.01
# For a larger or more complex task:
epsilon_start = 1.0
decay_rate = 0.999 # slower decay, more thorough exploration
epsilon_min = 0.01
# Rule of thumb:
# If training plateaus early with mediocre performance -> slow down decay
# If training converges slowly but policy looks good -> speed up decay
Tip: Print epsilon alongside your episode rewards so you can see the transition from exploration to exploitation. You should see reward increasing as epsilon falls — if it doesn’t, check your reward function and Q-update rule.
Run the epsilon decay calculation above with decay_rate = 0.99 and decay_rate = 0.999. At what episode does epsilon drop below 0.1 in each case? How does this affect the trade-off between exploration and learning speed?
Write a variant called softmax_action(q_table, state, temperature) that chooses actions with probability proportional to their Q-values (higher Q = higher probability, but lower Q actions still get chosen sometimes). The softmax probability of action a is exp(Q[s][a] / temperature) / sum(exp(Q[s][b] / temperature) for b in ACTIONS) — use import math and math.exp(). Compare its behaviour to ε-greedy. When might softmax be preferable?
Extend choose_action() to log every decision to a list: [(episode, step, state, action, mode)]. After training, count the ratio of EXPLORE to EXPLOIT decisions in the first 50 episodes vs the last 50 episodes.
“My agent is still mostly random after 200 episodes.”
Your decay rate is too slow. Try decay_rate = 0.99 (decay 1% per episode) instead of 0.999 (0.1% per episode).
“My agent converged to a policy quickly but it’s not very good.”
Your decay rate is too fast — the agent stopped exploring before it found better paths. Slow the decay down or increase epsilon_start.
“The ‘exploit’ path always picks the same action because all Q-values are equal.” This is expected early in training when everything is still 0. Add a tiny random tie-breaker:
return max(action_values, key=lambda a: action_values[a] + random.uniform(0, 1e-6))
This breaks ties randomly without meaningfully affecting the policy once values have diverged.
Next up: the full Q-learning update rule — combining everything we’ve built so far into the algorithm that drives learning.
You can use the arrows ← → on your keyboard to navigate between lessons.
Comments