Automatic Goal Generation for Reinforcement Learning Agents
arXiv:1705.06366
Abstract
Reinforcement learning is a powerful technique to train an agent to perform a task. However, an agent that is trained using reinforcement learning is only capable of achieving the single task that is specified via its reward function. Such an approach does not scale well to settings in which an agent needs to perform a diverse set of tasks, such as navigating to varying positions in a room or moving objects to varying locations. Instead, we propose a method that allows an agent to automatically discover the range of tasks that it is capable of performing. We use a generator network to propose tasks for the agent to try to achieve, specified as goal states. The generator network is optimized using adversarial training to produce tasks that are always at the appropriate level of difficulty for the agent. Our method thus automatically produces a curriculum of tasks for the agent to learn. We show that, by using this framework, an agent can efficiently and automatically learn to perform a wide set of tasks without requiring any prior knowledge of its environment. Our method can also learn to achieve tasks with sparse rewards, which traditionally pose significant challenges.
Accepted at ICML 2018, Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, PMLR 80, 2018
References in corpus (7)
- Active Learning of Inverse Models with Intrinsically Motivated Goal Exploration in Robots
- Hindsight Experience Replay
- Stochastic Neural Networks for Hierarchical Reinforcement Learning
- Learning and Transfer of Modulated Locomotor Controllers
- Surprise-Based Intrinsic Motivation for Deep Reinforcement Learning
- Automated Curriculum Learning for Neural Networks
- On the Effectiveness of Least Squares Generative Adversarial Networks
Cited by in corpus (37)
- Hindsight Experience Replay
- Episodic Curiosity through Reachability
- Reverse Curriculum Generation for Reinforcement Learning
- Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
- Efficient Exploration via State Marginal Matching
- GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms
- Unsupervised Meta-Learning for Reinforcement Learning
- An information-theoretic perspective on intrinsic motivation in reinforcement learning: a survey
- Embodied Visual Navigation with Automatic Curriculum Learning in Real Environments
- CURIOUS: Intrinsically Motivated Modular Multi-Goal Reinforcement Learning
- Reinforcement Learning for Robotic Manipulation using Simulated Locomotion Demonstrations
- Hindsight policy gradients
- Automatic Curriculum Learning through Value Disagreement
- Automated curricula through setter-solver interactions
- Many-Goals Reinforcement Learning
- Energy-Based Hindsight Experience Prioritization
- Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
- Recall Traces: Backtracking Models for Efficient Reinforcement Learning
- Rule Mining over Knowledge Graphs via Reinforcement Learning
- Towards Practical Multi-Object Manipulation using Relational Reinforcement Learning
- Hindsight Goal Ranking on Replay Buffer for Sparse Reward Environment
- SADA: Semantic Adversarial Diagnostic Attacks for Autonomous Applications
- Hyperparameter Auto-tuning in Self-Supervised Robotic Learning
- Learning to Plan Hierarchically from Curriculum
- Generating Automatic Curricula via Self-Supervised Active Domain Randomization
- Learning Deep Parameterized Skills from Demonstration for Re-targetable Visuomotor Control
- Weakly-Supervised Reinforcement Learning for Controllable Behavior
- HiER: Highlight Experience Replay for Boosting Off-Policy Reinforcement Learning Agents
- Competitive Experience Replay
- Hierarchical reinforcement learning for efficient exploration and transfer
- Learning Compositional Neural Programs for Continuous Control
- Hierarchical Representation Learning for Markov Decision Processes
- GloCAL: Glocalized Curriculum-Aided Learning of Multiple Tasks with Application to Robotic Grasping
- Stein Variational Goal Generation for adaptive Exploration in Multi-Goal Reinforcement Learning
- Meaning Versus Information, Prediction Versus Memory, and Question Versus Answer
- Region Growing Curriculum Generation for Reinforcement Learning
- Learning in Sparse Rewards settings through Quality-Diversity algorithms