KevsRobots Learning Platform
100% Percent Complete
By Kevin McAleer, 6 Minutes
Page last updated June 14, 2026

Letβs trace the full arc across both courses.
In the Reinforcement Learning for Beginners course, you started with a blank Python file and wrote everything yourself: the grid, the sensor, the reward function, the Q-table as a plain dictionary, the epsilon-greedy loop, and the Pico deployment code. That was deliberate β understanding every line is how the concepts stick.
In this course, you took those same ideas and moved them onto the tools the broader community uses. Here is what changed:
The environment interface:
state, reward, done = world.step(action) β a custom three-tuple, specific to your codeobs, reward, terminated, truncated, info = env.step(action) β the Gymnasium standard, compatible with any Gymnasium-aware libraryThe agent:
The deployment:
q_table.json loaded directly on the Piconn_policy.json in the same format, produced by querying the neural network over all 588 states β so the same Pico loader still worksThe physics did not change. The obstacles, the sensor, the rewards β all identical. What changed was the interface around them and the algorithm on top of them.
Work through this list. If anything feels uncertain, go back to the relevant lesson.
The Gymnasium API:
env.reset() returns (observation, info) and env.step(action) returns five valuesterminated means and how it differs from truncatedobservation_space and action_space are and how to inspect themenv.action_space.sample() to generate a random legal actionCustom environments:
gym.Env and implement reset() and step() correctlyobservation_space and action_space in __init__check_env and interpret its outputgym.register and create it with gym.makeWrappers and vectorisation:
RecordEpisodeStatistics adds to the info dictRewardWrapper that modifies rewards without editing the environmentSyncVectorEnv batches observations and rewards across N environmentsStable-Baselines3:
model.learn(total_timesteps=...)"MlpPolicy" means a fully-connected neural networkmodel.save() and load it with Algorithm.load()evaluate_policyDeployment:
The Q-table from the previous course was a dictionary with 588 keys and four float values per key. The neural network trained in this course is a function that takes a 4-dimensional input and produces four output values. Both are estimating the same thing: βhow much future reward should I expect if I take action A from state S?β
The difference is that the neural network can generalise across similar states, and it can handle observations β camera images, raw sensor arrays, joint angles β that no table could hold. For the 588-state BurgerBot world, both approaches work. For a real robot navigating a building, only the neural network approach scales.
You have learned Gymnasium, but Farama maintains a whole family of tools that extend it:
DQN implements both by default.If you want to revisit where you started:
The move from a 100-line hand-written environment to a Gymnasium class is mostly renaming and restructuring. The concepts β states, actions, rewards, episodes, exploration, exploitation β do not change. The standard API just makes them portable.
The move from a Q-table to a neural network is more significant. The network is harder to inspect, harder to debug, and harder to guarantee correct behaviour from. But it scales to problems the table never could.
Knowing both is a genuine superpower. You can reach for a Q-table when the state space is small and interpretability matters. You can reach for SB3 and PPO when the problem is too large for a table. And you now know how to wrap any environment β grid-world, physical robot, simulated arm β in a standard interface that either approach can use.
Good luck building.
You can use the arrows β β on your keyboard to navigate between lessons.
Comments