Implicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement Learning
arXiv:2010.14498
Abstract
We identify an implicit under-parameterization phenomenon in value-based deep RL methods that use bootstrapping: when value functions, approximated using deep neural networks, are trained with gradient descent using iterated regression onto target values generated by previous instances of the value network, more gradient updates decrease the expressivity of the current value network. We characterize this loss of expressivity via a drop in the rank of the learned value network features, and show that this typically corresponds to a performance drop. We demonstrate this phenomenon on Atari and Gym benchmarks, in both offline and online RL settings. We formally analyze this phenomenon and show that it results from a pathological interaction between bootstrapping and gradient-based optimization. We further show that mitigating implicit under-parameterization by controlling rank collapse can improve performance.
ICLR 2021. First two authors contributed equally. Website: https://agarwl.github.io/iup/
References in corpus (22)
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Conservative Q-Learning for Offline Reinforcement Learning
- Rainbow: Combining Improvements in Deep Reinforcement Learning
- Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
- Dopamine: A Research Framework for Deep Reinforcement Learning
- SBEED: Convergent Reinforcement Learning with Nonlinear Function Approximation
- On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
- Self-Distillation Amplifies Regularization in Hilbert Space
- Towards Characterizing Divergence in Deep Q-Learning
- Information-Theoretic Considerations in Batch Reinforcement Learning
- DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction
- Batch Value-function Approximation with Only Realizability
- Agnostic Q-learning with Function Approximation in Deterministic Systems: Tight Bounds on Approximation Error and Sample Complexity
- Harnessing Structures for Value-Based Planning and Reinforcement Learning
- A Finite-Time Analysis of Q-Learning with Neural Network Function Approximation
- Neural Temporal-Difference and Q-Learning Provably Converge to Global Optima
- Representations for Stable Off-Policy Reinforcement Learning
- The Utility of Sparse Representations for Control in Reinforcement Learning
- Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory
- On Catastrophic Interference in Atari 2600 Games
Cited by in corpus (5)
- Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
- D2RL: Deep Dense Architectures in Reinforcement Learning
- Spectral Normalisation for Deep Reinforcement Learning: an Optimisation Perspective
- The Difficulty of Passive Learning in Deep Reinforcement Learning
- Learning Off-Policy with Online Planning