Exploration vs Exploitation

Understand the ε-greedy strategy — when to trust what you know and when to try something new — and implement epsilon decay.

By Kevin McAleer,    6 Minutes

Page last updated June 13, 2026


Cover


Imagine you’re a robot with a half-learned Q-table. In your current state, turn_left looks promising (Q = 1.2) and forward looks risky (Q = -0.5). Should you always turn left?

Not necessarily. The Q-values are estimates based on experience so far — and early in training, there hasn’t been much experience. If you always pick the highest Q-value, you might be committing to a mediocre strategy before you’ve had a chance to discover better ones. But if you always pick randomly, you never benefit from what you’ve already learned.

This is the exploration-exploitation dilemma, and it sits at the heart of every reinforcement learning algorithm.


The ε-Greedy Strategy

The most popular solution is beautifully simple: most of the time, pick the best-known action (exploit); occasionally, pick a random action (explore).

ε (epsilon) is the probability of choosing a random action:

import random

def choose_action(q_table, state, epsilon):
    """
    ε-greedy action selection.

    With probability epsilon, choose a random action (explore).
    With probability 1-epsilon, choose the action with the
    highest Q-value (exploit).

    Args:
        q_table (dict): current Q-table (built using lesson 7's defaultdict)
        state (tuple): current state
        epsilon (float): exploration probability, between 0 and 1

    Returns:
        str: the chosen action
    """
    if random.random() < epsilon:
        # Explore: pick a uniformly random action
        return random.choice(ACTIONS)
    else:
        # Exploit: pick the action with the highest Q-value
        action_values = q_table[state]
        return max(action_values, key=lambda a: action_values[a])

With epsilon = 0.1, the agent explores 10% of the time and exploits its knowledge 90% of the time. This single parameter controls the entire balance.


Epsilon Decay: Getting Smarter Over Time

At the start of training, the Q-table is all zeros — the agent knows nothing. Exploration is valuable then because it generates diverse experience. But late in training, the Q-table holds hard-won knowledge. Exploring randomly at that point just wastes time.

Epsilon decay addresses this by gradually reducing epsilon over the course of training:

def decay_epsilon(epsilon, decay_rate, epsilon_min):
    """
    Reduce epsilon after each episode.

    Multiply epsilon by decay_rate (e.g. 0.995), but never let
    it fall below epsilon_min so the agent keeps a small amount
    of curiosity even when fully trained.

    Args:
        epsilon (float): current exploration probability
        decay_rate (float): multiplicative decay per episode (e.g. 0.995)
        epsilon_min (float): lowest allowed epsilon (e.g. 0.01)

    Returns:
        float: new (smaller) epsilon
    """
    return max(epsilon_min, epsilon * decay_rate)

# Typical training run
epsilon = 1.0        # start fully random (explore everything)
decay_rate = 0.995   # shrink epsilon by 0.5% each episode
epsilon_min = 0.01   # keep at least 1% exploration forever

for episode in range(500):
    # ... run episode using choose_action(q_table, state, epsilon) ...
    epsilon = decay_epsilon(epsilon, decay_rate, epsilon_min)

# After 500 episodes:
# epsilon ≈ 1.0 * (0.995 ** 500) ≈ 0.082
# Roughly 8% exploration by the end — still a little curious!

Let’s trace how epsilon evolves over time:

# Print epsilon at key checkpoints
epsilon = 1.0
decay_rate = 0.995
epsilon_min = 0.01
checkpoints = [0, 50, 100, 200, 300, 500, 750, 1000]
ep = 0
for episode in range(1001):
    if episode in checkpoints:
        print(f"Episode {episode:4d}: epsilon = {epsilon:.4f}")
    epsilon = max(epsilon_min, epsilon * decay_rate)

# Output:
# Episode    0: epsilon = 1.0000
# Episode   50: epsilon = 0.7783
# Episode  100: epsilon = 0.6058
# Episode  200: epsilon = 0.3670
# Episode  300: epsilon = 0.2222
# Episode  500: epsilon = 0.0816
# Episode  750: epsilon = 0.0236
# Episode 1000: epsilon = 0.0100  (clamped at epsilon_min)

The agent starts out almost completely random, gradually transitions to mostly exploiting its knowledge, and settles into a 99% exploit / 1% explore regime.


The Q-Table + ε-Greedy Working Together

Let’s trace through a few steps of the learning loop to see everything working together. We’ll use the Q-table from lesson 7:

from collections import defaultdict
import random

ACTIONS = ["forward", "turn_left", "turn_right", "stop"]

def make_q_table():
    return defaultdict(lambda: {action: 0.0 for action in ACTIONS})

def choose_action(q_table, state, epsilon):
    if random.random() < epsilon:
        return random.choice(ACTIONS)
    action_values = q_table[state]
    return max(action_values, key=lambda a: action_values[a])

# Set up
q_table = make_q_table()
epsilon = 0.8   # mostly exploring early on

# Simulate a few decisions
random.seed(42)
state = (1, 1, "east", "far")

for step in range(6):
    roll = random.random()
    if roll < epsilon:
        chosen = random.choice(ACTIONS)
        mode = "EXPLORE"
    else:
        chosen = max(q_table[state], key=lambda a: q_table[state][a])
        mode = "EXPLOIT"
    print(f"Step {step}: roll={roll:.3f} ({mode}) -> {chosen}")

Notice that when all Q-values are zero, exploit and explore give the same result (any action is equally “best”). The difference becomes meaningful once some episodes have run and the Q-values have diverged.


Tuning Epsilon Decay

The right decay rate depends on the task:

# For a small grid world (5x5, fast to explore):
epsilon_start = 1.0
decay_rate = 0.99      # faster decay, most exploration done in ~200 episodes
epsilon_min = 0.01

# For a larger or more complex task:
epsilon_start = 1.0
decay_rate = 0.999     # slower decay, more thorough exploration
epsilon_min = 0.01

# Rule of thumb:
# If training plateaus early with mediocre performance -> slow down decay
# If training converges slowly but policy looks good -> speed up decay

Tip: Print epsilon alongside your episode rewards so you can see the transition from exploration to exploitation. You should see reward increasing as epsilon falls — if it doesn’t, check your reward function and Q-update rule.


Try It Yourself

  1. Run the epsilon decay calculation above with decay_rate = 0.99 and decay_rate = 0.999. At what episode does epsilon drop below 0.1 in each case? How does this affect the trade-off between exploration and learning speed?

  2. Write a variant called softmax_action(q_table, state, temperature) that chooses actions with probability proportional to their Q-values (higher Q = higher probability, but lower Q actions still get chosen sometimes). The softmax probability of action a is exp(Q[s][a] / temperature) / sum(exp(Q[s][b] / temperature) for b in ACTIONS) — use import math and math.exp(). Compare its behaviour to ε-greedy. When might softmax be preferable?

  3. Extend choose_action() to log every decision to a list: [(episode, step, state, action, mode)]. After training, count the ratio of EXPLORE to EXPLOIT decisions in the first 50 episodes vs the last 50 episodes.


Common Issues

“My agent is still mostly random after 200 episodes.” Your decay rate is too slow. Try decay_rate = 0.99 (decay 1% per episode) instead of 0.999 (0.1% per episode).

“My agent converged to a policy quickly but it’s not very good.” Your decay rate is too fast — the agent stopped exploring before it found better paths. Slow the decay down or increase epsilon_start.

“The ‘exploit’ path always picks the same action because all Q-values are equal.” This is expected early in training when everything is still 0. Add a tiny random tie-breaker:

return max(action_values, key=lambda a: action_values[a] + random.uniform(0, 1e-6))

This breaks ties randomly without meaningfully affecting the policy once values have diverged.


Next up: the full Q-learning update rule — combining everything we’ve built so far into the algorithm that drives learning.


< Previous Next >

You can use the arrows  ← → on your keyboard to navigate between lessons.


What are you looking for?
Watch Videos Get Ideas Learn Something Read a Review Read the Blog Search
... Z Z z