Intrinsic Motivation and Automatic Curricula via Asymmetric Self-Play
arXiv:1703.05407
Abstract
We describe a simple scheme that allows an agent to learn about its environment in an unsupervised manner. Our scheme pits two versions of the same agent, Alice and Bob, against one another. Alice proposes a task for Bob to complete; and then Bob attempts to complete the task. In this work we will focus on two kinds of environments: (nearly) reversible environments and environments that can be reset. Alice will "propose" the task by doing a sequence of actions and then Bob must undo or repeat them, respectively. Via an appropriate reward structure, Alice and Bob automatically generate a curriculum of exploration, enabling unsupervised training of the agent. When Bob is deployed on an RL task within the environment, this unsupervised training reduces the number of supervised episodes needed to learn, and in some cases converges to a higher reward.
Published in ICLR 2018
References in corpus (4)
Cited by in corpus (47)
- Dota 2 with Large Scale Deep Reinforcement Learning
- Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
- Parameter Space Noise for Exploration
- Hindsight Experience Replay
- Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
- Emergent Complexity via Multi-Agent Competition
- Automatic Goal Generation for Reinforcement Learning Agents
- Reverse Curriculum Generation for Reinforcement Learning
- Illuminating Generalization in Deep Reinforcement Learning through Procedural Level Generation
- Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
- Universal Planning Networks
- Unsupervised Meta-Learning for Reinforcement Learning
- Learning with AMIGo: Adversarially Motivated Intrinsic Goals
- Embodied Visual Navigation with Automatic Curriculum Learning in Real Environments
- A social path to human-like artificial intelligence
- Re-evaluating Evaluation
- Reinforcement Learning for Robotic Manipulation using Simulated Locomotion Demonstrations
- Unicorn: Continual Learning with a Universal, Off-policy Agent
- Automatic Curriculum Learning through Value Disagreement
- Automated curricula through setter-solver interactions
- Energy-Based Hindsight Experience Prioritization
- Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
- Rule Mining over Knowledge Graphs via Reinforcement Learning
- Recall Traces: Backtracking Models for Efficient Reinforcement Learning
- Learning Latent Plans from Play
- ScreenerNet: Learning Self-Paced Curriculum for Deep Neural Networks
- Deep Intrinsically Motivated Continuous Actor-Critic for Efficient Robotic Visuomotor Skill Learning
- Recent Advances in Neural Program Synthesis
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering Skills
- SMiRL: Surprise Minimizing Reinforcement Learning in Unstable Environments
- Ecological Reinforcement Learning
- Learning to Prove Theorems by Learning to Generate Theorems
- Is Curiosity All You Need? On the Utility of Emergent Behaviours from Curious Exploration
- Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied Navigation
- Adversarial Environment Generation for Learning to Navigate the Web
- Reinforcement Learning with Success Induced Task Prioritization
- HiER: Highlight Experience Replay for Boosting Off-Policy Reinforcement Learning Agents
- Curriculum By Smoothing
- Competitive Experience Replay
- Amplifying the Imitation Effect for Reinforcement Learning of UCAV's Mission Execution
- Switching Isotropic and Directional Exploration with Parameter Space Noise in Deep Reinforcement Learning
- Stein Variational Goal Generation for adaptive Exploration in Multi-Goal Reinforcement Learning
- Region Growing Curriculum Generation for Reinforcement Learning
- Planning in Dynamic Environments with Conditional Autoregressive Models
- Population Based Training for Data Augmentation and Regularization in Speech Recognition
- Explore and Control with Adversarial Surprise
- Growing Action Spaces