Selected work

Reinforcement learning · simulation

DiffDriveRL

DiffDriveRL is my simulator for training a differential-drive robot to reach waypoints. The policy controls the left and right wheels in a 144-inch-square field.

Policy
PPO · two wheel commands
Environment
144-inch-square simulated field
Training
Curriculum with obstacles and waypoints
Hardware results
Not tested on a physical robot

Where I started

The simulator gives the policy a target and two wheel commands.

I built a PPO environment for a differential-drive robot. Each action sets one command per wheel. The observation includes the target’s direction and distance, the robot’s heading error, and route hints.

The project includes scripts to train a policy, replay a run, and compare a candidate checkpoint with a baseline.

What I built

A policy that controls two wheels.

I added obstacles, blocked routes, recovery starts, and a final-heading requirement to the waypoint task.

Training stages

Train on open routes, then add obstacles.

The profiles move from clear routes to light obstacles, full obstacles, and connected waypoints. Some profiles also sample blocked routes, reverse targets, and recovery starts.

Rewards

Reward progress and penalize rough control.

The environment rewards progress and alignment, then adds penalties for collisions, spinning in place, and jittery actions. Those are simulator reward choices, not measurements of a real robot.

Evaluation

Compare a candidate before promoting it.

The evaluation scripts run episodes from a checkpoint. A separate gate compares a candidate with a baseline; the code only promotes a checkpoint when the gate is deliberately used and passes.

Choices I made

Increase route difficulty over time.

  • Start with open routes.

    The curriculum can add obstacles after the policy sees the simpler waypoint task.

  • Include routes the robot cannot drive straight through.

    Blocked-route examples make the policy practice moving around an obstacle toward the goal.

  • Keep evaluation separate from training.

    The candidate gate checks a saved policy against a baseline before replacement.