Deep Reinforcement Learning and the Deadly Triad
arXiv:1812.02648
Abstract
We know from reinforcement learning theory that temporal difference learning can fail in certain cases. Sutton and Barto (2018) identify a deadly triad of function approximation, bootstrapping, and off-policy learning. When these three properties are combined, learning can diverge with the value estimates becoming unbounded. However, several algorithms successfully combine these three properties, which indicates that there is at least a partial gap in our understanding. In this work, we investigate the impact of the deadly triad in practice, in the context of a family of popular deep reinforcement learning models - deep Q-networks trained with experience replay - analysing how the components of this system play a role in the emergence of the deadly triad, and in the agent's performance
References in corpus (3)
Cited by in corpus (39)
- Bootstrap your own latent: A new approach to self-supervised Learning
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Behavior Regularized Offline Reinforcement Learning
- MOReL : Model-Based Offline Reinforcement Learning
- Information-Theoretic Considerations in Batch Reinforcement Learning
- Uncovering Instabilities in Variational-Quantum Deep Q-Networks
- A Geometric Perspective on Optimal Representations for Reinforcement Learning
- Learning to Reach Goals via Iterated Supervised Learning
- Shared Experience Actor-Critic for Multi-Agent Reinforcement Learning
- Generalizable Episodic Memory for Deep Reinforcement Learning
- deep-significance - Easy and Meaningful Statistical Significance Testing in the Age of Neural Networks
- Deep Q-learning: a robust control approach
- Training Agents using Upside-Down Reinforcement Learning
- Lipschitzness Is All You Need To Tame Off-policy Generative Adversarial Imitation Learning
- On the Estimation Bias in Double Q-Learning
- Training Larger Networks for Deep Reinforcement Learning
- Symbolic Equation Solving via Reinforcement Learning
- Towards Understanding Cooperative Multi-Agent Q-Learning with Value Factorization
- Online Hyper-parameter Tuning in Off-policy Learning via Evolutionary Strategies
- Breaking the Deadly Triad with a Target Network
- Convergent and Efficient Deep Q Network Algorithm
- GRAC: Self-Guided and Self-Regularized Actor-Critic
- Self-supervised Graph Representation Learning via Bootstrapping
- Qgraph-bounded Q-learning: Stabilizing Model-Free Off-Policy Deep Reinforcement Learning
- PBCS : Efficient Exploration and Exploitation Using a Synergy between Reinforcement Learning and Motion Planning
- PER-ETD: A Polynomially Efficient Emphatic Temporal Difference Learning Method
- Approximating two value functions instead of one: towards characterizing a new family of Deep Reinforcement Learning algorithms
- Fuzzy Tiling Activations: A Simple Approach to Learning Sparse Representations Online
- Ensemble Bootstrapping for Q-Learning
- Off-policy Reinforcement Learning with Optimistic Exploration and Distribution Correction
- Policy Gradients Incorporating the Future
- The Difficulty of Passive Learning in Deep Reinforcement Learning
- Reducing Conservativeness Oriented Offline Reinforcement Learning
- Analyzing Reinforcement Learning Benchmarks with Random Weight Guessing
- Evolution of Q Values for Deep Q Learning in Stable Baselines
- Uncertainty-aware Low-Rank Q-Matrix Estimation for Deep Reinforcement Learning
- MBCAL: Sample Efficient and Variance Reduced Reinforcement Learning for Recommender Systems
- Hybrid BYOL-ViT: Efficient approach to deal with small datasets