Torque-Driven RL for Quadruped Locomotion

A torque-driven RL framework for the Unitree B1 quadruped achieves 3.5 m/s speeds and stair climbing without exteroceptive sensors, using NVIDIA Isaac Lab.

axonn bots
axonn bots
·2 min read
A torque-driven reinforcement learning framework for the Unitree B1 quadruped achieves speeds of 3.5 m/s and can traverse stairs without exteroceptive sensors. The work, published at IEEE/SICE SII 2026, demonstrates the potential of torque control for heavyweight quadrupeds.

Reinforcement learning (RL) for legged robots has demonstrated impressive adaptability to new and challenging terrain. However, traditional RL locomotion frameworks are position-based, making policies less adaptable and requiring state estimation techniques like linear velocity in the observation space. Moreover, these frameworks often use small, lightweight quadrupeds that are limited in their viability for high-complexity tasks.[reference:44]

A new paper, "Towards Torque-Driven Reinforcement Learning for Quadruped Locomotion," explores an RL torque control framework for heavyweight high-torque quadrupeds.[reference:45] The framework can traverse rough terrain and effectively track a desired linear velocity without requiring knowledge of the agent's current velocity.[reference:46]

Using NVIDIA's Isaac Sim and Isaac Lab, simulation results of the RL torque control policy are shown on the Unitree B1 quadruped, achieving speeds of 3.5 m/s and 1.5 rad/s.[reference:47] In addition, the quadruped can walk up and down stairs without the aid of an exteroceptive sensor.[reference:48]

The paper, authored by Jordan Dowdy and one other researcher, was published in the 2026 IEEE/SICE International Symposium on System Integration (SII), pp. 1259-1264.[reference:49][reference:50]

This work represents a shift toward torque-level control for RL-based locomotion, offering improved adaptability and performance on heavier robots. The ability to traverse stairs without exteroceptive sensors is particularly noteworthy, as it reduces the reliance on expensive and complex perception systems.