MADE: Exploration via Maximizing Deviation from Explored Regions
arXiv:2106.10268
Abstract
In online reinforcement learning (RL), efficient exploration remains particularly challenging in high-dimensional environments with sparse rewards. In low-dimensional environments, where tabular parameterization is possible, count-based upper confidence bound (UCB) exploration methods achieve minimax near-optimal rates. However, it remains unclear how to efficiently implement UCB in realistic RL tasks that involve non-linear function approximation. To address this, we propose a new exploration approach via \textit{maximizing} the deviation of the occupancy of the next policy from the explored regions. We add this term as an adaptive regularizer to the standard RL objective to balance exploration vs. exploitation. We pair the new objective with a provably convergent algorithm, giving rise to a new intrinsic reward that adjusts existing bonuses. The proposed intrinsic reward is easy to implement and combine with other existing RL algorithms to conduct exploration. As a proof of concept, we evaluate the new intrinsic reward on tabular examples across a variety of model-based and model-free algorithms, showing improvements over count-only exploration strategies. When tested on navigation and locomotion tasks from MiniGrid and DeepMind Control Suite benchmarks, our approach significantly improves sample efficiency over state-of-the-art methods. Our code is available at https://github.com/tianjunz/MADE.
28 pages, 10 figures
References in corpus (49)
- Continuous control with deep reinforcement learning
- DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
- Large-Scale Study of Curiosity-Driven Learning
- #Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning
- Variational Information Maximisation for Intrinsically Motivated Reinforcement Learning
- Exploration by Random Network Distillation
- Reinforcement Learning with Augmented Data
- Bridging the Gap Between Value and Policy Based Reinforcement Learning
- Provably Efficient Reinforcement Learning with Linear Function Approximation
- Go-Explore: a New Approach for Hard-Exploration Problems
- dm_control: Software and Tasks for Continuous Control
- Variational Intrinsic Control
- Deep Exploration via Randomized Value Functions
- Contextual Decision Processes with Low Bellman Rank are PAC-Learnable
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels
- On the Theory of Policy Gradient Methods: Optimality, Approximation, and Distribution Shift
- A unified view of entropy-regularized Markov decision processes
- Provably Efficient Maximum Entropy Exploration
- A Theory of Regularized Markov Decision Processes
- Efficient Exploration via State Marginal Matching
- Model-Based Active Exploration
- Never Give Up: Learning Directed Exploration Strategies
- Model-based RL in Contextual Decision Processes: PAC bounds and Exponential Improvements over Model-free Approaches
- Optimism in Reinforcement Learning with Generalized Linear Function Approximation
- Provably efficient RL with Rich Observations via Latent State Decoding
- Learning with AMIGo: Adversarially Motivated Intrinsic Goals
- Reinforcement Learning with Prototypical Representations
- Fast Global Convergence of Natural Policy Gradient Methods with Entropy Regularization
- Approximate Policy Iteration Schemes: A Comparison
- FLAMBE: Structural Complexity and Representation Learning of Low Rank MDPs
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments
- Fast active learning for pure exploration in reinforcement learning
- Is Reinforcement Learning More Difficult Than Bandits? A Near-optimal Algorithm Escaping the Curse of Horizon
- EMI: Exploration with Mutual Information
- Provably Efficient Reinforcement Learning for Discounted MDPs with Feature Mapping
- PC-PG: Policy Cover Directed Exploration for Provable Policy Gradient Learning
- State Entropy Maximization with Random Encoders for Efficient Exploration
- Variance Reduction Methods for Sublinear Reinforcement Learning
- Scheduled Intrinsic Drive: A Hierarchical Take on Intrinsically Motivated Exploration
- BeBold: Exploration Beyond the Boundary of Explored Regions
- Worst-Case Regret Bounds for Exploration via Randomized Value Functions
- UCB Momentum Q-learning: Correcting the bias without forgetting
- Nearly Minimax Optimal Reinforcement Learning for Discounted MDPs
- Marginalized State Distribution Entropy Regularization in Policy Optimization
- Regret Bounds for Discounted MDPs
- Cautiously Optimistic Policy Optimization and Exploration with Linear Function Approximation
- Provable Model-based Nonlinear Bandit and Reinforcement Learning: Shelve Optimism, Embrace Virtual Curvature
- Provably Correct Optimization and Exploration with Non-linear Policies