Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal Transformers
arXiv:2107.03996
Abstract
We propose to address quadrupedal locomotion tasks using Reinforcement Learning (RL) with a Transformer-based model that learns to combine proprioceptive information and high-dimensional depth sensor inputs. While learning-based locomotion has made great advances using RL, most methods still rely on domain randomization for training blind agents that generalize to challenging terrains. Our key insight is that proprioceptive states only offer contact measurements for immediate reaction, whereas an agent equipped with visual sensory observations can learn to proactively maneuver environments with obstacles and uneven terrain by anticipating changes in the environment many steps ahead. In this paper, we introduce LocoTransformer, an end-to-end RL method that leverages both proprioceptive states and visual observations for locomotion control. We evaluate our method in challenging simulated environments with different obstacles and uneven terrain. We transfer our learned policy from simulation to a real robot by running it indoors and in the wild with unseen obstacles and terrain. Our method not only significantly improves over baselines, but also achieves far better generalization performance, especially when transferred to the real robot. Our project page with videos is at https://rchalyang.github.io/LocoTransformer/ .
Our project page with videos is at https://RchalYang.github.io/LocoTransformer
References in corpus (15)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Language Models are Few-Shot Learners
- Learning agile and dynamic motor skills for legged robots
- Emergence of Locomotion Behaviours in Rich Environments
- Generating Long Sequences with Sparse Transformers
- Learning to Navigate in Complex Environments
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Reinforcement Learning with Unsupervised Auxiliary Tasks
- Decoupling Representation Learning from Reinforcement Learning
- Policies Modulating Trajectory Generators
- Learning a Contact-Adaptive Controller for Robust, Efficient Legged Locomotion
- From Pixels to Legs: Hierarchical Learning of Quadruped Locomotion
- RMA: Rapid Motor Adaptation for Legged Robots
- Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers
- GLiDE: Generalizable Quadrupedal Locomotion in Diverse Environments with a Centroidal Model
Cited by in corpus (4)
- Robust High-speed Running for Quadruped Robots via Deep Reinforcement Learning
- Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
- Multimodal Information Bottleneck for Deep Reinforcement Learning with Multiple Sensors
- Vision-Guided Quadrupedal Locomotion in the Wild with Multi-Modal Delay Randomization