Incentivizing Exploration In Reinforcement Learning With Deep Predictive Models
arXiv:1507.00814
Abstract
Achieving efficient and scalable exploration in complex domains poses a major challenge in reinforcement learning. While Bayesian and PAC-MDP approaches to the exploration problem offer strong formal guarantees, they are often impractical in higher dimensions due to their reliance on enumerating the state-action space. Hence, exploration in complex domains is often performed with simple epsilon-greedy methods. In this paper, we consider the challenging Atari games domain, which requires processing raw pixel inputs and delayed rewards. We evaluate several more sophisticated exploration strategies, including Thompson sampling and Boltzman exploration, and propose a new exploration method based on assigning exploration bonuses from a concurrently learned model of the system dynamics. By parameterizing our learned model with a neural network, we are able to develop a scalable and efficient approach to exploration bonuses that can be applied to tasks with complex, high-dimensional state spaces. In the Atari domain, our method provides the most consistent improvement across a range of games that pose a major challenge for prior methods. In addition to raw game-scores, we also develop an AUC-100 metric for the Atari Learning domain to evaluate the impact of exploration on this benchmark.
References in corpus (1)
Cited by in corpus (130)
- A Brief Survey of Deep Reinforcement Learning
- Prioritized Experience Replay
- Dueling Network Architectures for Deep Reinforcement Learning
- An Introduction to Deep Reinforcement Learning
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation
- Deep Reinforcement Learning: An Overview
- Exploration in Deep Reinforcement Learning: A Survey
- VIME: Variational Information Maximizing Exploration
- Deep Exploration via Bootstrapped DQN
- Emergent Tool Use From Multi-Agent Autocurricula
- Parameter Space Noise for Exploration
- Large-Scale Study of Curiosity-Driven Learning
- #Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning
- Exploration by Random Network Distillation
- Unifying Count-Based Exploration and Intrinsic Motivation
- Go-Explore: a New Approach for Hard-Exploration Problems
- Episodic Curiosity through Reachability
- Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
- Meta-Reinforcement Learning of Structured Exploration Strategies
- Control of Memory, Active Perception, and Action in Minecraft
- Reinforcement Learning in Healthcare: A Survey
- Improving Exploration in Evolution Strategies for Deep Reinforcement Learning via a Population of Novelty-Seeking Agents
- Surprise-Based Intrinsic Motivation for Deep Reinforcement Learning
- Multi-Objective Deep Reinforcement Learning
- A survey on intrinsic motivation in reinforcement learning
- Efficient Exploration via State Marginal Matching
- Deep Successor Reinforcement Learning
- Dialog-based Language Learning
- Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks
- Never Give Up: Learning Directed Exploration Strategies
- Self-Consistent Trajectory Autoencoder: Hierarchical Reinforcement Learning with Trajectory Embeddings
- On Improving Deep Reinforcement Learning for POMDPs
- Dynamics-Aware Unsupervised Discovery of Skills
- Some Considerations on Learning to Explore via Meta-Reinforcement Learning
- Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control
- Skew-Fit: State-Covering Self-Supervised Reinforcement Learning
- Unsupervised Meta-Learning for Reinforcement Learning
- Estimating Risk and Uncertainty in Deep Reinforcement Learning
- Deep Reinforcement Learning in Parameterized Action Space
- Learning with AMIGo: Adversarially Motivated Intrinsic Goals
- Learning Gentle Object Manipulation with Curiosity-Driven Deep Reinforcement Learning
- Value Prediction Network
- Curiosity-Driven Experience Prioritization via Density Estimation
- Towards a Common Implementation of Reinforcement Learning for Multiple Robotic Tasks
- Learning Exploration Policies for Navigation
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments
- On Learning Intrinsic Rewards for Policy Gradient Methods
- Learning Multimodal Transition Dynamics for Model-Based Reinforcement Learning
- Learning to Explore with Meta-Policy Gradient
- Learning Self-Imitating Diverse Policies
- Benchmarking Bonus-Based Exploration Methods on the Arcade Learning Environment
- Human-Level Reinforcement Learning through Theory-Based Modeling, Exploration, and Planning
- Exploratory Gradient Boosting for Reinforcement Learning in Complex Domains
- Deep Intrinsically Motivated Continuous Actor-Critic for Efficient Robotic Visuomotor Skill Learning
- State Entropy Maximization with Random Encoders for Efficient Exploration
- Curiosity in exploring chemical space: Intrinsic rewards for deep molecular reinforcement learning
- Learning latent state representation for speeding up exploration
- BeBold: Exploration Beyond the Boundary of Explored Regions
- See, Hear, Explore: Curiosity via Audio-Visual Association
- Sequence Modeling of Temporal Credit Assignment for Episodic Reinforcement Learning
- Learning Dynamics Model in Reinforcement Learning by Incorporating the Long Term Future
- Efficient exploration with Double Uncertain Value Networks
- Randomized Value Functions via Multiplicative Normalizing Flows
- Variational Bayesian Reinforcement Learning with Regret Bounds
- DAQN: Deep Auto-encoder and Q-Network
- Adaptive trajectory-constrained exploration strategy for deep reinforcement learning
- World Discovery Models
- Exploration and preference satisfaction trade-off in reward-free learning
- Adapting Behaviour via Intrinsic Reward: A Survey and Empirical Study
- Deep Learning for Reward Design to Improve Monte Carlo Tree Search in ATARI Games
- Never Forget: Balancing Exploration and Exploitation via Learning Optical Flow
- Learning-Driven Exploration for Reinforcement Learning
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven Exploration
- Long-Term Visitation Value for Deep Exploration in Sparse Reward Reinforcement Learning
- Novelty Search in Representational Space for Sample Efficient Exploration
- Autonomous exploration for navigating in non-stationary CMPs
- Model-Based Regularization for Deep Reinforcement Learning with Transcoder Networks
- Active World Model Learning with Progress Curiosity
- A Survey of Exploration Methods in Reinforcement Learning
- Non-local Policy Optimization via Diversity-regularized Collaborative Exploration
- Rank the Episodes: A Simple Approach for Exploration in Procedurally-Generated Environments
- Hashing over Predicted Future Frames for Informed Exploration of Deep Reinforcement Learning
- Intrinsic Motivation for Encouraging Synergistic Behavior
- Reinforcement Learning with Probabilistically Complete Exploration
- Information Maximizing Exploration with a Latent Dynamics Model
- Competitive Experience Replay
- Multimodal Reward Shaping for Efficient Exploration in Reinforcement Learning
- Empowerment-driven Exploration using Mutual Information Estimation
- PBCS : Efficient Exploration and Exploitation Using a Synergy between Reinforcement Learning and Motion Planning
- Exploration in Approximate Hyper-State Space for Meta Reinforcement Learning
- Online reinforcement learning with sparse rewards through an active inference capsule
- AutoEG: Automated Experience Grafting for Off-Policy Deep Reinforcement Learning
- Sample Efficient Reinforcement Learning via Model-Ensemble Exploration and Exploitation
- Energy-based Surprise Minimization for Multi-Agent Value Factorization
- Amplifying the Imitation Effect for Reinforcement Learning of UCAV's Mission Execution
- Beyond Exponentially Discounted Sum: Automatic Learning of Return Function
- Goal-oriented Trajectories for Efficient Exploration
- Adversarial Active Exploration for Inverse Dynamics Model Learning
- Touch-based Curiosity for Sparse-Reward Tasks
- Disentangling Controllable Object through Video Prediction Improves Visual Reinforcement Learning
- Near-Optimal Reward-Free Exploration for Linear Mixture MDPs with Plug-in Solver
- Density-based Curriculum for Multi-goal Reinforcement Learning with Sparse Rewards
- Micro-Objective Learning : Accelerating Deep Reinforcement Learning through the Discovery of Continuous Subgoals
- ISL: A novel approach for deep exploration
- Deep Reinforcement Learning with Weighted Q-Learning
- Regret Bounds and Reinforcement Learning Exploration of EXP-based Algorithms
- Learning to Plan Optimistically: Uncertainty-Guided Deep Exploration via Latent Model Ensembles
- Scaling All-Goals Updates in Reinforcement Learning Using Convolutional Neural Networks
- Off-policy Reinforcement Learning with Optimistic Exploration and Distribution Correction
- GAN-based Intrinsic Exploration For Sample Efficient Reinforcement Learning
- Don't Do What Doesn't Matter: Intrinsic Motivation with Action Usefulness
- Towards Robust Bisimulation Metric Learning
- State-Aware Variational Thompson Sampling for Deep Q-Networks
- Reannealing of Decaying Exploration Based On Heuristic Measure in Deep Q-Network
- Clustered Reinforcement Learning
- DeepFoldit -- A Deep Reinforcement Learning Neural Network Folding Proteins
- Unbiased Deep Reinforcement Learning: A General Training Framework for Existing and Future Algorithms
- Exploring More When It Needs in Deep Reinforcement Learning
- Knowledge is reward: Learning optimal exploration by predictive reward cashing
- Neural Embedding for Physical Manipulations
- MIME: Mutual Information Minimisation Exploration
- Influence-Based Reinforcement Learning for Intrinsically-Motivated Agents
- Region Growing Curriculum Generation for Reinforcement Learning
- UAV-assisted Online Machine Learning over Multi-Tiered Networks: A Hierarchical Nested Personalized Federated Learning Approach
- Learning to Shape Rewards using a Game of Two Partners
- Decoupled Exploration and Exploitation Policies for Sample-Efficient Reinforcement Learning
- Learning Good Representation via Continuous Attention
- Perturbation-based exploration methods in deep reinforcement learning
- Learning Efficient and Effective Exploration Policies with Counterfactual Meta Policy
- Intrinsic Exploration as Multi-Objective RL