KevsRobots Learning Platform
12% Percent Complete
By Kevin McAleer, 5 Minutes
Page last updated June 13, 2026

Let’s start with an honest conversation about the robots most of us have built.
You write a forward() function. You write a turn_left(). You carefully measure the obstacle threshold — “if the sensor reads less than 20 cm, turn right”. You test it on your kitchen table. It works! You put it on the floor of the living room. It crashes into the sofa leg. You tweak the threshold to 25 cm. Now it spins in the corner for thirty seconds before eventually escaping. You stay up until midnight adding more if statements.
Sound familiar?
The typical beginner robot loop looks something like this:
def navigate():
while True:
distance = get_distance()
if distance < 20:
stop()
turn_right()
sleep(0.4)
else:
forward()
sleep(0.1)
This is perfectly reasonable code. The logic is clear and it’s easy to reason about. But it contains a hidden assumption: the programmer already knows the right response to every situation the robot will encounter.
When the world is simple and predictable, that is fine. When the world is messy — different floor surfaces, objects of varying heights, variable lighting, a curious cat — the assumption breaks down fast.
Every new situation requires a new rule. Every new rule might conflict with an existing one. The codebase grows, becomes brittle, and starts hiding bugs that only appear at 2 am on a Tuesday.
Think about how you’d train a dog to sit.
You don’t hand the dog a rulebook. You don’t reprogram its brain. You create a feedback signal: when the dog sits, something good happens (a treat). When the dog does something you don’t want, the treat doesn’t come, or a gentle “no” signals disapproval.
Over many repetitions, the dog figures out the strategy on its own. The behaviour emerges from feedback, not from instructions.
Reinforcement learning is exactly this idea, applied to software:
The robot’s code does not contain the answer at the start. The answer is discovered through experience.
Here’s the key insight: in a static, fully known environment, hardcoded rules can be optimal. In a dynamic environment — one where the layout changes, sensor readings are noisy, or the task evolves — a fixed rule set will always be playing catch-up.
A learned policy doesn’t just execute a script. It evaluates the current situation and chooses the action that has worked best in the past for situations that looked like this one. When you train it in a varied enough environment, it generalises gracefully to situations it hasn’t seen before.
This is a genuine paradigm shift. Let’s be concrete about the difference:
| Hardcoded approach | Learned approach |
|---|---|
| You write the rules | The robot discovers the rules |
| Fails in new situations | Generalises to new situations |
| Debugging is adding more if-statements | Debugging is adjusting rewards |
| Fast to implement for simple tasks | Takes longer to set up, scales further |
| Transparent: you can read the logic | Interpretable via the Q-table |
For simple, well-defined environments, hardcoded logic is often the right choice — don’t reach for RL when a if distance < 20: turn_right() genuinely solves the problem. But once the environment gets complicated enough that you’re managing dozens of rules and still seeing failures, it’s worth learning a better tool.
By lesson 10 of this course, you’ll have a Python program that does something like this:
# No hardcoded rules about what to do in each situation.
# The robot discovers the best action by trying many thousands
# of moves in a simulated grid world and receiving rewards.
for episode in range(500):
state = env.reset()
total_reward = 0
while not done:
action = choose_action(q_table, state, epsilon)
next_state, reward, done = env.step(action)
q_table = update_q(q_table, state, action, reward, next_state)
state = next_state
total_reward += reward
print(f"Episode {episode}: total reward = {total_reward:.1f}")
Each episode, the robot tries to navigate a grid, gets rewards for good moves and penalties for collisions, and updates a table of “how good is each action in each situation?” values. After enough episodes, that table holds a surprisingly competent policy — no rules written by hand.
Before moving on, have a think about your own robot projects:
These questions don’t have right or wrong answers yet — we’re just starting to think in RL terms. Keep your answers in mind as we work through the next few lessons.
Next up, we’ll give all these ideas precise names and pin down exactly what “agent”, “environment”, “state”, and “action” mean in a robotics context.
You can use the arrows ← → on your keyboard to navigate between lessons.
Comments