Reinforcement learning · simulation
DiffDriveRL
DiffDriveRL is my simulator for training a differential-drive robot to reach waypoints. The policy controls the left and right wheels in a 144-inch-square field.
- Policy
- PPO · two wheel commands
- Environment
- 144-inch-square simulated field
- Training
- Curriculum with obstacles and waypoints
- Hardware results
- Not tested on a physical robot
Where I started
The simulator gives the policy a target and two wheel commands.
I built a PPO environment for a differential-drive robot. Each action sets one command per wheel. The observation includes the target’s direction and distance, the robot’s heading error, and route hints.
The project includes scripts to train a policy, replay a run, and compare a candidate checkpoint with a baseline.
What I built
A policy that controls two wheels.
I added obstacles, blocked routes, recovery starts, and a final-heading requirement to the waypoint task.
Training stages
Train on open routes, then add obstacles.
The profiles move from clear routes to light obstacles, full obstacles, and connected waypoints. Some profiles also sample blocked routes, reverse targets, and recovery starts.
Rewards
Reward progress and penalize rough control.
The environment rewards progress and alignment, then adds penalties for collisions, spinning in place, and jittery actions. Those are simulator reward choices, not measurements of a real robot.
Evaluation
Compare a candidate before promoting it.
The evaluation scripts run episodes from a checkpoint. A separate gate compares a candidate with a baseline; the code only promotes a checkpoint when the gate is deliberately used and passes.
Choices I made
Increase route difficulty over time.
Start with open routes.
The curriculum can add obstacles after the policy sees the simpler waypoint task.
Include routes the robot cannot drive straight through.
Blocked-route examples make the policy practice moving around an obstacle toward the goal.
Keep evaluation separate from training.
The candidate gate checks a saved policy against a baseline before replacement.