KevsRobots Learning Platform
6% Percent Complete
By Kevin McAleer, 4 Minutes
Page last updated June 14, 2026

In the Reinforcement Learning for Beginners course you built everything by hand: a custom GridWorld class, a Q-table stored in a plain Python dictionary, an epsilon-greedy action selector, and a reward function tuned for the BurgerBot maze. That was exactly the right place to start. Understanding every line of a hand-rolled environment is how the concepts stick.
This course takes those same ideas and ports them onto the tools that real researchers and industry teams use every day. Gymnasium (maintained by the Farama Foundation) is the standard API for describing environments. Stable-Baselines3 (SB3) is a library of battle-tested agents β DQN, PPO, A2C, and more β that plug straight into any Gymnasium environment. Together they let you swap algorithms, share code with the community, and eventually scale up to problems where a lookup table simply runs out of road.
The through-line project is the BurgerBot grid-world you already know. By the end you will have wrapped it as a first-class Gymnasium environment, validated it with the official checker, trained a neural-network policy on it with SB3, and exported that policy as a JSON table that the Pico loader from the previous course can read directly.
This course picks up where the Reinforcement Learning for Beginners course finishes. Before starting here, make sure you have completed:
env.step() and why terminated and truncated are separateDiscrete, Box, MultiDiscreteenv.close()gym.Env contract and how to write a custom environment from scratchTimeLimit, RecordEpisodeStatistics, and custom wrappersAfter completing this course, you will be able to:
terminated and truncated and why it matters for traininggym.Env subclass that passes check_envThis course runs entirely on a regular computer. You do not need a Pico, a BurgerBot, or any other hardware for the lessons themselves. The export lesson produces a file you can later load onto a Pico using the loader from the previous course.
pip install gymnasium stable-baselines3 β both install via pip; stable-baselines3 pulls in PyTorch automaticallyA GPU is not required. All training in this course runs comfortably on CPU in a few minutes.
Code blocks show complete, runnable scripts. Every script in this course was tested with Gymnasium 1.x and Stable-Baselines3 2.x before publication. Where training takes a long time, the lesson notes a reduced total_timesteps value you can use to verify the script runs, then suggests a larger value for a better-trained policy.
Notes and warnings appear as blockquotes:
Note: Extra context that is useful but not essential.
Warning: Something that will bite you if you skip it.
Each lesson that introduces runnable code ends with a βTry It Yourselfβ section and a βCommon Issuesβ section. Work through the exercises β the fastest way to build intuition is to break things on purpose and read the error messages.
You can use the arrows β β on your keyboard to navigate between lessons.
Comments