Count-Based Exploration with Neural Density Models
arXiv:1703.01310
Abstract
Bellemare et al. (2016) introduced the notion of a pseudo-count, derived from a density model, to generalize count-based exploration to non-tabular reinforcement learning. This pseudo-count was used to generate an exploration bonus for a DQN agent and combined with a mixed Monte Carlo update was sufficient to achieve state of the art on the Atari 2600 game Montezuma's Revenge. We consider two questions left open by their work: First, how important is the quality of the density model for exploration? Second, what role does the Monte Carlo update play in exploration? We answer the first question by demonstrating the use of PixelCNN, an advanced neural density model for images, to supply a pseudo-count. In particular, we examine the intrinsic difficulties in adapting Bellemare et al.'s approach when assumptions about the model are violated. The result is a more practical and general algorithm requiring no special apparatus. We combine PixelCNN pseudo-counts with different agent architectures to dramatically improve the state of the art on several hard Atari games. One surprising finding is that the mixed Monte Carlo update is a powerful facilitator of exploration in the sparsest of settings, including Montezuma's Revenge.
References in corpus (1)
Cited by in corpus (96)
- An Introduction to Deep Reinforcement Learning
- Generative Modeling by Estimating Gradients of the Data Distribution
- Noisy Networks for Exploration
- Exploration in Deep Reinforcement Learning: A Survey
- Emergent Tool Use From Multi-Agent Autocurricula
- Parameter Space Noise for Exploration
- Large-Scale Study of Curiosity-Driven Learning
- Hindsight Experience Replay
- Exploration by Random Network Distillation
- Agent57: Outperforming the Atari Human Benchmark
- Go-Explore: a New Approach for Hard-Exploration Problems
- Episodic Curiosity through Reachability
- Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
- Evolved Policy Gradients
- Improving Exploration in Evolution Strategies for Deep Reinforcement Learning via a Population of Novelty-Seeking Agents
- Feature Control as Intrinsic Motivation for Hierarchical Reinforcement Learning
- Learning Montezuma's Revenge from a Single Demonstration
- Some Considerations on Learning to Explore via Meta-Reinforcement Learning
- Contingency-Aware Exploration in Reinforcement Learning
- The Uncertainty Bellman Equation and Exploration
- Rule-Based Reinforcement Learning for Efficient Robot Navigation with Space Reduction
- An information-theoretic perspective on intrinsic motivation in reinforcement learning: a survey
- Counting to Explore and Generalize in Text-based Games
- Learning with AMIGo: Adversarially Motivated Intrinsic Goals
- The NetHack Learning Environment
- InfoBot: Transfer and Exploration via the Information Bottleneck
- Reinforcement Learning with Prototypical Representations
- Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation
- Re-evaluating Evaluation
- Coordinated Exploration via Intrinsic Rewards for Multi-Agent Reinforcement Learning
- Computational Theories of Curiosity-Driven Learning
- Q-Learning in enormous action spaces via amortized approximate maximization
- Recall Traces: Backtracking Models for Efficient Reinforcement Learning
- Focus on Impact: Indoor Exploration with Intrinsic Motivation
- Sample-efficient Reinforcement Learning Representation Learning with Curiosity Contrastive Forward Dynamics Model
- Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement Learning
- Benchmarking Bonus-Based Exploration Methods on the Arcade Learning Environment
- State Entropy Maximization with Random Encoders for Efficient Exploration
- Landmark-Guided Subgoal Generation in Hierarchical Reinforcement Learning
- Is Deep Reinforcement Learning Really Superhuman on Atari? Leveling the playing field
- Scheduled Intrinsic Drive: A Hierarchical Take on Intrinsically Motivated Exploration
- BeBold: Exploration Beyond the Boundary of Explored Regions
- Deep Curiosity Search: Intra-Life Exploration Can Improve Performance on Challenging Deep Reinforcement Learning Problems
- Regularized Softmax Deep Multi-Agent -Learning
- Harnessing Structures for Value-Based Planning and Reinforcement Learning
- Boredom-driven curious learning by Homeo-Heterostatic Value Gradients
- Variational Bayesian Reinforcement Learning with Regret Bounds
- Surprising Negative Results for Generative Adversarial Tree Search
- World Discovery Models
- Behavior From the Void: Unsupervised Active Pre-Training
- Adaptive Reward-Free Exploration
- Long-Term Visitation Value for Deep Exploration in Sparse Reward Reinforcement Learning
- Novelty Search in Representational Space for Sample Efficient Exploration
- Deep Abstract Q-Networks
- Approximate Exploration through State Abstraction
- Temporal Difference Uncertainties as a Signal for Exploration
- Locally Persistent Exploration in Continuous Control Tasks with Sparse Rewards
- Continuous Doubly Constrained Batch Reinforcement Learning
- Diversity-Enriched Option-Critic
- Hashing over Predicted Future Frames for Informed Exploration of Deep Reinforcement Learning
- Non-local Policy Optimization via Diversity-regularized Collaborative Exploration
- What Can Learned Intrinsic Rewards Capture?
- Flow-based Intrinsic Curiosity Module
- Implicit Generative Modeling for Efficient Exploration
- Learning What to Memorize: Using Intrinsic Motivation to Form Useful Memory in Partially Observable Reinforcement Learning
- Review, Analysis and Design of a Comprehensive Deep Reinforcement Learning Framework
- Exploration in Approximate Hyper-State Space for Meta Reinforcement Learning
- Behavior-Guided Actor-Critic: Improving Exploration via Learning Policy Behavior Representation for Deep Reinforcement Learning
- Momentum in Reinforcement Learning
- Novel Policy Seeking with Constrained Optimization
- Amplifying the Imitation Effect for Reinforcement Learning of UCAV's Mission Execution
- REMAX: Relational Representation for Multi-Agent Exploration
- Guided Exploration with Proximal Policy Optimization using a Single Demonstration
- Accelerating Reinforcement Learning with a Directional-Gaussian-Smoothing Evolution Strategy
- Exploration by Maximizing Rényi Entropy for Reward-Free RL Framework
- Directed Exploration in PAC Model-Free Reinforcement Learning
- Adversarial Active Exploration for Inverse Dynamics Model Learning
- GAN-based Intrinsic Exploration For Sample Efficient Reinforcement Learning
- Regularly Updated Deterministic Policy Gradient Algorithm
- Learning to Plan Optimistically: Uncertainty-Guided Deep Exploration via Latent Model Ensembles
- ISL: A novel approach for deep exploration
- Exploring Unknown States with Action Balance
- Importance Sampling based Exploration in Q Learning
- Combine PPO with NES to Improve Exploration
- Planning with Exploration: Addressing Dynamics Bottleneck in Model-based Reinforcement Learning
- Decoupled Exploration and Exploitation Policies for Sample-Efficient Reinforcement Learning
- Neural Embedding for Physical Manipulations
- Generative Exploration and Exploitation
- Learning to Shape Rewards using a Game of Two Partners
- When should agents explore?
- Count-Based Temperature Scheduling for Maximum Entropy Reinforcement Learning
- Perturbation-based exploration methods in deep reinforcement learning
- ACDER: Augmented Curiosity-Driven Experience Replay
- Region Growing Curriculum Generation for Reinforcement Learning
- Being curious about the answers to questions: novelty search with learned attention
- On the potential for open-endedness in neural networks