Emergence of Locomotion Behaviours in Rich Environments
arXiv:1707.02286
Abstract
The reinforcement learning paradigm allows, in principle, for complex behaviours to be learned directly from simple reward signals. In practice, however, it is common to carefully hand-design the reward function to encourage a particular solution, or to derive it from demonstration data. In this paper explore how a rich environment can help to promote the learning of complex behavior. Specifically, we train agents in diverse environmental contexts, and find that this encourages the emergence of robust behaviours that perform well across a suite of tasks. We demonstrate this principle for locomotion -- behaviours that are known for their sensitivity to the choice of reward. We train several simulated bodies on a diverse set of challenging terrains and obstacles, using a simple reward function based on forward progress. Using a novel scalable variant of policy gradient reinforcement learning, our agents learn to run, jump, crouch and turn as required by the environment without explicit reward-based guidance. A visual depiction of highlights of the learned behavior can be viewed following https://youtu.be/hx_bgoTF7bs .
References in corpus (3)
Cited by in corpus (119)
- A Brief Survey of Deep Reinforcement Learning
- Learning agile and dynamic motor skills for legged robots
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
- Embodied Intelligence via Learning and Evolution
- Character Controllers Using Motion VAEs
- Benchmarking Model-Based Reinforcement Learning
- dm_control: Software and Tasks for Continuous Control
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- Applying deep reinforcement learning to active flow control in turbulent conditions
- Autonomous Unmanned Aerial Vehicle Navigation using Reinforcement Learning: A Systematic Review
- AutoSlim: Towards One-Shot Architecture Search for Channel Numbers
- Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions
- Non-Smooth Newton Methods for Deformable Multi-Body Dynamics
- Exploring Model-based Planning with Policy Networks
- Measuring and modeling the motor system with machine learning
- Iterative Reinforcement Learning Based Design of Dynamic Locomotion Skills for Cassie
- Enhanced POET: Open-Ended Reinforcement Learning through Unbounded Invention of Learning Challenges and their Solutions
- Transfer in Deep Reinforcement Learning Using Successor Features and Generalised Policy Improvement
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames
- Neural Graph Evolution: Towards Efficient Automatic Robot Design
- Quantum imaginary time evolution steered by reinforcement learning
- Self-supervised Learning of Image Embedding for Continuous Control
- Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal Transformers
- One Policy to Control Them All: Shared Modular Policies for Agent-Agnostic Control
- TriFinger: An Open-Source Robot for Learning Dexterity
- Generalization in Dexterous Manipulation via Geometry-Aware Multi-Task Learning
- Reinforcement Learning with Competitive Ensembles of Information-Constrained Primitives
- Whole-Body Control of a Mobile Manipulator using End-to-End Reinforcement Learning
- Strategies for Using Proximal Policy Optimization in Mobile Puzzle Games
- Reactive Stepping for Humanoid Robots using Reinforcement Learning: Application to Standing Push Recovery on the Exoskeleton Atalante
- Sequence Modeling of Temporal Credit Assignment for Episodic Reinforcement Learning
- Modelling Generalized Forces with Reinforcement Learning for Sim-to-Real Transfer
- Reward-Conditioned Policies
- Generation of ice states through deep reinforcement learning
- Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
- Behavior Priors for Efficient Reinforcement Learning
- Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves
- Human-Inspired Multi-Agent Navigation using Knowledge Distillation
- Development of collective behavior in newborn artificial agents
- Asynchronous Methods for Model-Based Reinforcement Learning
- Hedging using reinforcement learning: Contextual -Armed Bandit versus -learning
- Evolving the Behavior of Machines: From Micro to Macroevolution
- Ecological Reinforcement Learning
- Evaluating Agents without Rewards
- Physics-Based Dexterous Manipulations with Estimated Hand Poses and Residual Reinforcement Learning
- Towards General and Autonomous Learning of Core Skills: A Case Study in Locomotion
- Deep neuroethology of a virtual rodent
- A Deep Reinforcement Learning Framework for Eco-driving in Connected and Automated Hybrid Electric Vehicles
- AI-MOLE: Autonomous Iterative Motion Learning for Unknown Nonlinear Dynamics with Extensive Experimental Validation
- The Benchmark Lottery
- Offline Meta-Reinforcement Learning with Online Self-Supervision
- Curriculum in Gradient-Based Meta-Reinforcement Learning
- Zero-Shot Terrain Generalization for Visual Locomotion Policies
- Fast and Efficient Locomotion via Learned Gait Transitions
- Guided Curriculum Learning for Walking Over Complex Terrain
- Synthesizing Long-Term 3D Human Motion and Interaction in 3D Scenes
- Run, skeleton, run: skeletal model in a physics-based simulation
- Distributed Heuristic Multi-Agent Path Finding with Communication
- Cascade Attribute Learning Network
- Physically Embedded Planning Problems: New Challenges for Reinforcement Learning
- Marathon Environments: Multi-Agent Continuous Control Benchmarks in a Modern Video Game Engine
- Reinforcement Learning with Adaptive Curriculum Dynamics Randomization for Fault-Tolerant Robot Control
- Imagined Value Gradients: Model-Based Policy Optimization with Transferable Latent Dynamics Models
- Deep Reinforcement Learning in Fluid Mechanics: a promising method for both Active Flow Control and Shape Optimization
- Combining Benefits from Trajectory Optimization and Deep Reinforcement Learning
- Dynamics and Domain Randomized Gait Modulation with Bezier Curves for Sim-to-Real Legged Locomotion
- ALLSTEPS: Curriculum-driven Learning of Stepping Stone Skills
- Robust Domain Randomised Reinforcement Learning through Peer-to-Peer Distillation
- Model-free and Bayesian Ensembling Model-based Deep Reinforcement Learning for Particle Accelerator Control Demonstrated on the FERMI FEL
- Balance Between Efficient and Effective Learning: Dense2Sparse Reward Shaping for Robot Manipulation with Environment Uncertainty
- Thalamocortical motor circuit insights for more robust hierarchical control of complex sequences
- Learning Generalizable Locomotion Skills with Hierarchical Reinforcement Learning
- HJB Optimal Feedback Control with Deep Differential Value Functions and Action Constraints
- Bottom-Up Meta-Policy Search
- Competitive Experience Replay
- An Actor-Critic-Attention Mechanism for Deep Reinforcement Learning in Multi-view Environments
- Simple Sensor Intentions for Exploration
- Explicitly Encouraging Low Fractional Dimensional Trajectories Via Reinforcement Learning
- Cross-Domain Imitation Learning via Optimal Transport
- Augmented Random Search for Quadcopter Control: An alternative to Reinforcement Learning
- Snowflake: Scaling GNNs to High-Dimensional Continuous Control via Parameter Freezing
- Beyond Tabula-Rasa: a Modular Reinforcement Learning Approach for Physically Embedded 3D Sokoban
- Competitiveness of MAP-Elites against Proximal Policy Optimization on locomotion tasks in deterministic simulations
- Adaptive Smoothing Path Integral Control
- Learning 6DoF Grasping Using Reward-Consistent Demonstration
- Distributed Deep Reinforcement Learning: An Overview
- SeRO: Self-Supervised Reinforcement Learning for Recovery from Out-of-Distribution Situations
- Learning Diverse Policies with Soft Self-Generated Guidance
- From Recognition to Prediction: Analysis of Human Action and Trajectory Prediction in Video
- Partially Connected Automated Vehicle Cooperative Control Strategy with a Deep Reinforcement Learning Approach
- Multi-intersection Traffic Optimisation: A Benchmark Dataset and a Strong Baseline
- Discovering Diverse Athletic Jumping Strategies
- Discovering Generalizable Skills via Automated Generation of Diverse Tasks
- Mission schedule of agile satellites based on Proximal Policy Optimization Algorithm
- Evaluating model-based planning and planner amortization for continuous control
- Design of AoI-Aware 5G Uplink Scheduler UsingReinforcement Learning
- Hierarchical Skills for Efficient Exploration
- Learning Agile Locomotion Skills with a Mentor
- SURREAL-System: Fully-Integrated Stack for Distributed Deep Reinforcement Learning
- Metric-Based Imitation Learning Between Two Dissimilar Anthropomorphic Robotic Arms
- High-Level Perceptual Similarity is Enabled by Learning Diverse Tasks
- Critic PI2: Master Continuous Planning via Policy Improvement with Path Integrals and Deep Actor-Critic Reinforcement Learning
- Zero Shot Learning on Simulated Robots
- Efficient Reinforcement Learning Development with RLzoo
- Training Transition Policies via Distribution Matching for Complex Tasks
- Homotopy Based Reinforcement Learning with Maximum Entropy for Autonomous Air Combat
- Containerized Distributed Value-Based Multi-Agent Reinforcement Learning
- Vision-Guided Quadrupedal Locomotion in the Wild with Multi-Modal Delay Randomization
- Direct Random Search for Fine Tuning of Deep Reinforcement Learning Policies
- Quadruped Locomotion on Non-Rigid Terrain using Reinforcement Learning
- Lifetime policy reuse and the importance of task capacity
- Learning a Skill-sequence-dependent Policy for Long-horizon Manipulation Tasks
- Adaptive Policy Transfer in Reinforcement Learning
- Parallelized Reverse Curriculum Generation
- Learning and Exploring Motor Skills with Spacetime Bounds
- Efficient Hyperparameter Optimization for Physics-based Character Animation
- Recognition and Synthesis of Object Transport Motion
- Learning walk and trot from the same objective using different types of exploration