Exploration in Deep Reinforcement Learning: From Single-Agent to Multiagent Domain
arXiv:2109.06668 · doi:10.1109/TNNLS.2023.3236361
Abstract
Deep Reinforcement Learning (DRL) and Deep Multi-agent Reinforcement Learning (MARL) have achieved significant successes across a wide range of domains, including game AI, autonomous vehicles, robotics, and so on. However, DRL and deep MARL agents are widely known to be sample inefficient that millions of interactions are usually needed even for relatively simple problem settings, thus preventing the wide application and deployment in real-industry scenarios. One bottleneck challenge behind is the well-known exploration problem, i.e., how efficiently exploring the environment and collecting informative experiences that could benefit policy learning towards the optimal ones. This problem becomes more challenging in complex environments with sparse rewards, noisy distractions, long horizons, and non-stationary co-learners. In this paper, we conduct a comprehensive survey on existing exploration methods for both single-agent and multi-agent RL. We start the survey by identifying several key challenges to efficient exploration. Beyond the above two main branches, we also include other notable exploration methods with different ideas and techniques. In addition to algorithmic analysis, we provide a comprehensive and unified empirical comparison of different exploration methods for DRL on a set of commonly used benchmarks. According to our algorithmic and empirical investigation, we finally summarize the open problems of exploration in DRL and deep MARL and point out a few future directions.
Accepted by IEEE Transactions on Neural Networks and Learning Systems (TNNLS)
References in corpus (26)
- A Brief Survey of Deep Reinforcement Learning
- Exploration in Deep Reinforcement Learning: A Survey
- Parameter Space Noise for Exploration
- Safe Exploration in Continuous Action Spaces
- Variational Intrinsic Control
- Trial without Error: Towards Safe Reinforcement Learning via Human Intervention
- Bootstrapped Thompson Sampling and Deep Exploration
- Episodic Multi-agent Reinforcement Learning with Curiosity-Driven Exploration
- Reinforcement Learning with Prototypical Representations
- Conservative Safety Critics for Exploration
- Cooperative Exploration for Multi-Agent Deep Reinforcement Learning
- A New Framework for Multi-Agent Reinforcement Learning -- Centralized Training and Exploration with Decentralized Execution via Policy Distillation
- Variational Dynamic for Self-Supervised Exploration in Deep Reinforcement Learning
- Learning latent state representation for speeding up exploration
- Scheduled Intrinsic Drive: A Hierarchical Take on Intrinsically Motivated Exploration
- Exploration with Unreliable Intrinsic Reward in Multi-Agent Reinforcement Learning
- Dynamic Bottleneck for Robust Self-Supervised Exploration
- Towards Safe Reinforcement Learning with a Safety Editor Policy
- DFAC Framework: Factorizing the Value Function via Quantile Mixture for Multi-Agent Distributional Q-Learning
- Directed Exploration for Reinforcement Learning
- Self-Imitation Advantage Learning
- MULEX: Disentangling Exploitation from Exploration in Deep RL
- Ready Policy One: World Building Through Active Learning
- Principled Exploration via Optimistic Bootstrapping and Backward Induction
- Co-Imitation Learning without Expert Demonstration
- Efficient exploration of zero-sum stochastic games
Cited by in corpus (3)
- GenSafe: A Generalizable Safety Enhancer for Safe Reinforcement Learning Algorithms Based on Reduced Order Markov Decision Process Model
- Inverse RL Scene Dynamics Learning for Nonlinear Predictive Control in Autonomous Vehicles
- Meta-Offline and Distributional Multi-Agent RL for Risk-Aware Decision-Making