A simulated grey quadruped robot standing on a dark checkered floor

RoboticsPart 2 of 2Drawing 3.2 of 37

Learning to walk

A policy I've trained in simulation since August 2026. This part covers what it can do, the three things that broke getting there, and why none of it has been on the robot.

Sheet 02 of 04/Walking

Teaching it to walk

The robot's apart on a shelf right now, so everything below comes from simulation, with 64 robots running at once.

Verification criteria / simulation only
CriterionBarResultPolicy
Walks forward5.0 m6.7 mFlat ground
Top speednone set0.72 m/sFlat ground
Holds a straight line7 deg4.5 degFlat ground
Stays upright100 of 100Flat ground
Climbs a 10 deg slope98 of 100 uprightSlope
Walks sidewaysshort by 3 mm/sFlat ground
Criteria passed1110Neither alone

One policy trained on flat ground made most of this table, and a second one trained on slopes made the climb. No single policy has done both yet.

A simulated quadruped robot walking forward across a blue checkered plane, the camera trailing it
PL 09Run 182 at 7,896 iterations, driven forward at 0.49 m/s. Across eight robots it covers 2.2 to 3.3 m in eight seconds, and none of them fall. This is the newest policy, but it isn't the best one.

Sheet 03 of 04/What broke

What I got wrong

I built the software stack on a bad CAD export, so every policy trained on it was learning to walk as a robot that didn't exist. In August 2026 I archived all of it to a git tag and started over from SolidWorks.

Sheet 04 of 04/Next

Where it goes next

Right now the policy trots, so the body falls between steps and the servos catch it twice a stride. I don't think hobby servos can take that on the real robot, so this version stops here. The learning setup moves on to Gray 2.0.