Shifat Santo

stampede

Thousands of simulated bipeds share one policy and learn to walk together on a single GPU.

Sep 2026Research code, private

430kcontrol steps per second, 4,096 robots on one RTX 3090

A small biped, the dodo, built on Open Duck Mini v2 legs, is trained with PPO across thousands of copies at once on batched MuJoCo Warp. A browser viewer shows whole waves of learners at the same time: early ones collapse, later ones shuffle, the best ones stride.

The speed comes from capturing the physics substep loop as a CUDA graph. Only the simulator step lives inside the graph; the training state stays outside, and per-robot torque limits are rewritten in place so the graph is captured once. On a rented RTX 3090 that is 430k control steps per second at 4,096 robots, against 15k on a laptop.

Two findings changed what the results mean. The orientation input the policy saw was the body axis in the world frame, something no real IMU can measure, 25 to 41 degrees away from what a sensor would report. Swapping in a batched Madgwick filter fixed it, and it invalidated every earlier checkpoint for transfer. Later, the endurance evaluator turned out never to count falls, so every earlier "walks for 20 minutes" was false. Re-run correctly, the median walk before a fall is 0.3 to 0.7 minutes for every policy.

What does not work yet: no physical robot walks, policies fall on balance within a minute, and getting up after a fall is still at zero.