EX2: Exploration with Exemplar Models for Deep Reinforcement Learning
arXiv:1703.01260
Abstract
Deep reinforcement learning algorithms have been shown to learn complex tasks using highly general policy classes. However, sparse reward problems remain a significant challenge. Exploration methods based on novelty detection have been particularly successful in such settings but typically require generative or predictive models of the observations, which can be difficult to train when the observations are very high-dimensional and complex, as in the case of raw images. We propose a novelty detection algorithm for exploration that is based entirely on discriminatively trained exemplar models, where classifiers are trained to discriminate each visited state against all others. Intuitively, novel states are easier to distinguish against other states seen during training. We show that this kind of discriminative modeling corresponds to implicit density estimation, and that it can be combined with count-based exploration to produce competitive results on a range of popular benchmark tasks, including state-of-the-art results on challenging egocentric observations in the vizDoom benchmark.
Cited by in corpus (33)
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- Large-Scale Study of Curiosity-Driven Learning
- Agent57: Outperforming the Atari Human Benchmark
- Episodic Curiosity through Reachability
- WAIC, but Why? Generative Ensembles for Robust Anomaly Detection
- Reinforcement Learning in Healthcare: A Survey
- A survey on intrinsic motivation in reinforcement learning
- UCB Exploration via Q-Ensembles
- Dynamics-Aware Unsupervised Discovery of Skills
- Residual Policy Learning
- An information-theoretic perspective on intrinsic motivation in reinforcement learning: a survey
- Learning Exploration Policies for Navigation
- Perfect density models cannot guarantee anomaly detection
- Learning Self-Imitating Diverse Policies
- Learning latent state representation for speeding up exploration
- Self-Supervised Exploration via Disagreement
- Improving Robot Dual-System Motor Learning with Intrinsically Motivated Meta-Control and Latent-Space Experience Imagination
- Unsupervised Out-of-Distribution Detection with Batch Normalization
- Meta-learning curiosity algorithms
- Optimistic Exploration even with a Pessimistic Initialisation
- Safety Augmented Value Estimation from Demonstrations (SAVED): Safe Deep Model-Based RL for Sparse Cost Robotic Tasks
- Exploring Exploration: Comparing Children with RL Agents in Unified Environments
- Locally Persistent Exploration in Continuous Control Tasks with Sparse Rewards
- A Survey of Exploration Methods in Reinforcement Learning
- Multi-Path Policy Optimization
- Competitive Experience Replay
- Diversity-Driven Extensible Hierarchical Reinforcement Learning
- Empowerment-driven Exploration using Mutual Information Estimation
- Curiosity creates Diversity in Policy Search
- Clustered Reinforcement Learning
- Optimising Stochastic Routing for Taxi Fleets with Model Enhanced Reinforcement Learning
- Biased Estimates of Advantages over Path Ensembles
- Neural Embedding for Physical Manipulations