Surprise-Based Intrinsic Motivation for Deep Reinforcement Learning
arXiv:1703.01732
Abstract
Exploration in complex domains is a key challenge in reinforcement learning, especially for tasks with very sparse rewards. Recent successes in deep reinforcement learning have been achieved mostly using simple heuristic exploration strategies such as -greedy action selection or Gaussian control noise, but there are many tasks where these methods are insufficient to make any learning progress. Here, we consider more complex heuristics: efficient and scalable exploration strategies that maximize a notion of an agent's surprise about its experiences via intrinsic motivation. We propose to learn a model of the MDP transition probabilities concurrently with the policy, and to form intrinsic rewards that approximate the KL-divergence of the true transition probabilities from the learned model. One of our approximations results in using surprisal as intrinsic motivation, while the other gives the -step learning progress. We show that our incentives enable agents to succeed in a wide range of environments with high-dimensional state spaces and very sparse rewards, including continuous control tasks and games in the Atari RAM domain, outperforming several other heuristic exploration techniques.
Appeared in Deep RL Workshop at NIPS 2016
References in corpus (2)
Cited by in corpus (15)
- Parameter Space Noise for Exploration
- Planning to Explore via Self-Supervised World Models
- Learning Gentle Object Manipulation with Curiosity-Driven Deep Reinforcement Learning
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments
- Self-Supervised Exploration via Disagreement
- World Discovery Models
- Bayesian Curiosity for Efficient Exploration in Reinforcement Learning
- Autonomous exploration for navigating in non-stationary CMPs
- Non-local Policy Optimization via Diversity-regularized Collaborative Exploration
- Deep Reinforcement Learning in Fluid Mechanics: a promising method for both Active Flow Control and Shape Optimization
- Active World Model Learning with Progress Curiosity
- Efficient Exploration through Intrinsic Motivation Learning for Unsupervised Subgoal Discovery in Model-Free Hierarchical Reinforcement Learning
- MIME: Mutual Information Minimisation Exploration
- Neural Embedding for Physical Manipulations
- Learning Efficient and Effective Exploration Policies with Counterfactual Meta Policy