Reverse Curriculum Generation for Reinforcement Learning
arXiv:1707.05300
Abstract
Many relevant tasks require an agent to reach a certain state, or to manipulate objects into a desired configuration. For example, we might want a robot to align and assemble a gear onto an axle or insert and turn a key in a lock. These goal-oriented tasks present a considerable challenge for reinforcement learning, since their natural reward function is sparse and prohibitive amounts of exploration are required to reach the goal and receive some learning signal. Past approaches tackle these problems by exploiting expert demonstrations or by manually designing a task-specific reward shaping function to guide the learning agent. Instead, we propose a method to learn these tasks without requiring any prior knowledge other than obtaining a single state in which the task is achieved. The robot is trained in reverse, gradually learning to reach the goal from a set of start states increasingly far from the goal. Our method automatically generates a curriculum of start states that adapts to the agent's performance, leading to efficient training on goal-oriented tasks. We demonstrate our approach on difficult simulated navigation and fine-grained manipulation problems, not solvable by state-of-the-art reinforcement learning methods.
Published at the 1st Conference on Robot Learning (CoRL 2017)
References in corpus (6)
- Active Learning of Inverse Models with Intrinsically Motivated Goal Exploration in Robots
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
- Automatic Goal Generation for Reinforcement Learning Agents
- Intrinsic Motivation and Automatic Curricula via Asymmetric Self-Play
- Automated Curriculum Learning for Neural Networks
- Towards Generalization and Simplicity in Continuous Control
Cited by in corpus (52)
- Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research
- Intrinsic Motivation and Automatic Curricula via Asymmetric Self-Play
- Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation
- Multi-task learning for natural language processing in the 2020s: where are we going?
- Learning Montezuma's Revenge from a Single Demonstration
- Curiosity Driven Exploration of Learned Disentangled Goal Spaces
- Meta reinforcement learning as task inference
- Unsupervised Learning of Goal Spaces for Intrinsically Motivated Goal Exploration
- Learning with AMIGo: Adversarially Motivated Intrinsic Goals
- Learning to Locomote: Understanding How Environment Design Matters for Deep Reinforcement Learning
- Unicorn: Continual Learning with a Universal, Off-policy Agent
- Automatic Curriculum Learning through Value Disagreement
- Computational Theories of Curiosity-Driven Learning
- Energy-Based Hindsight Experience Prioritization
- Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
- Accuracy-based Curriculum Learning in Deep Reinforcement Learning
- Backplay: "Man muss immer umkehren"
- Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning
- BeBold: Exploration Beyond the Boundary of Explored Regions
- Forward-Backward Reinforcement Learning
- Variational Automatic Curriculum Learning for Sparse-Reward Cooperative Multi-Agent Problems
- Time Reversal as Self-Supervision
- Learning Goal Embeddings via Self-Play for Hierarchical Reinforcement Learning
- Time Limits in Reinforcement Learning
- Curiosity-Driven Multi-Criteria Hindsight Experience Replay
- Self-supervised Learning of Distance Functions for Goal-Conditioned Reinforcement Learning
- Self-Paced Contextual Reinforcement Learning
- Curriculum in Gradient-Based Meta-Reinforcement Learning
- Learning Dense Rewards for Contact-Rich Manipulation Tasks
- Cascade Attribute Learning Network
- Training Stronger Baselines for Learning to Optimize
- ALLSTEPS: Curriculum-driven Learning of Stepping Stone Skills
- Curriculum By Smoothing
- Multiple-objective Reinforcement Learning for Inverse Design and Identification
- Learning to Navigate the Web
- PBCS : Efficient Exploration and Exploitation Using a Synergy between Reinforcement Learning and Motion Planning
- Coarse-to-Fine Curriculum Learning
- Interaction-limited Inverse Reinforcement Learning
- Competitive Experience Replay
- A Deep Reinforcement Learning Architecture for Multi-stage Optimal Control
- Hierarchical deep reinforcement learning controlled three-dimensional navigation of microrobots in blood vessels
- Cascade Attribute Network: Decomposing Reinforcement Learning Control Policies using Hierarchical Neural Networks
- Episodic Self-Imitation Learning with Hindsight
- Policy Gradients Incorporating the Future
- Expanding Motor Skills through Relay Neural Networks
- Growing Action Spaces
- A Probabilistic Interpretation of Self-Paced Learning with Applications to Reinforcement Learning
- Parallelized Reverse Curriculum Generation
- Region Growing Curriculum Generation for Reinforcement Learning
- Deep Reinforcement Learning for Complex Manipulation Tasks with Sparse Feedback
- Micro/Nano Motor Navigation and Localization via Deep Reinforcement Learning
- Lucid Dreaming for Experience Replay: Refreshing Past States with the Current Policy