1 paper
Roy Xing, Seyoung Ree, Brian Plancher
Reinforcement learning (RL) for locomotion frequently converges to locally optimal but undeployable behaviors, such as vibrating limbs or scooting on the torso, that maximize retur…