KevsRobots Learning Platform
100% Percent Complete
By Kevin McAleer, 6 Minutes
Page last updated June 13, 2026

You made it! Letβs take a moment to look back at the journey from βmy robot crashes into the sofaβ to βmy robot has a trained policy it discovered on its own.β
Over the course of 14 lessons, you:
defaultdict with JSON serialisation for Pico deploymentUse this as a revision guide. You should be able to explain each item in your own words:
Core vocabulary:
The framework:
The algorithm:
Deployment:
Beyond tabular:
This is the one equation worth memorising. Here it is one more time as commented Python:
# Q-learning update β commit this to memory
# Bellman-derived, Watkins 1989, still the heart of modern RL
q_table[state][action] += alpha * (
reward # what I just got
+ gamma * max(q_table[next_state].values()) # discounted best future
- q_table[state][action] # minus my current estimate (TD error)
)
# alpha (learning rate): how boldly I update β typically 0.1β0.3
# gamma (discount): how much I trust the future β typically 0.9β0.99
# TD error: positive = I was too pessimistic, negative = too optimistic
Youβve covered tabular Q-learning end-to-end. Here are the natural next steps, in roughly increasing order of difficulty:
MicroPython Robotics Projects If you havenβt completed this course, it covers the hardware foundations (motors, sensors, PWM) that the deployment lessons assumed. A great place to consolidate the physical side.
Learn ROS The Robot Operating System is the industry-standard framework for building complex multi-component robots. If you want to scale from a two-wheeled Pico robot to a full robot arm or mobile platform, ROS is the next major skill to acquire.
NumPy β before any deep learning framework, get comfortable with numerical arrays. The Pandas and NumPy course on this site is a good starting point.
OpenAI Gymnasium β the standard RL environment API. Try the FrozenLake, CartPole, and MountainCar environments with your Q-learning code first, then graduate to DQN for CartPole.
PyTorch or JAX β the deep learning frameworks used in most modern RL research. PyTorch has excellent beginner tutorials and an active community.
Spinning Up in Deep RL β OpenAIβs free educational resource: spinningup.openai.com. Covers policy gradients, PPO, SAC, and more with clean implementations.
When you started this course, a robot that learns might have seemed like something reserved for research labs with expensive GPUs and machine-learning PhDs. I hope this course has shown you that the core ideas are genuinely accessible β a Python dictionary, a reward signal, and enough patience to let a virtual robot crash into walls several hundred times is all you really need to get started.
The same Bellman equation that powers cutting-edge robotics research is running in your q_table[state][action] += alpha * td_error line. The scale is different. The principle is identical.
Keep building. Keep experimenting. And keep letting your robots learn.
Thank you for working through this course. If youβd like to share what you built, post it on social media and tag @kevinmcaleer28 β Iβd love to see your trained robots in action.
You can use the arrows β β on your keyboard to navigate between lessons.
Comments