DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction
arXiv:2003.07305
Abstract
Deep reinforcement learning can learn effective policies for a wide range of tasks, but is notoriously difficult to use due to instability and sensitivity to hyperparameters. The reasons for this remain unclear. When using standard supervised methods (e.g., for bandits), on-policy data collection provides "hard negatives" that correct the model in precisely those states and actions that the policy is likely to visit. We call this phenomenon "corrective feedback." We show that bootstrapping-based Q-learning algorithms do not necessarily benefit from this corrective feedback, and training on the experience collected by the algorithm is not sufficient to correct errors in the Q-function. In fact, Q-learning and related methods can exhibit pathological interactions between the distribution of experience collected by the agent and the policy induced by training on that experience, leading to potential instability, sub-optimal convergence, and poor results when learning from noisy, sparse or delayed rewards. We demonstrate the existence of this problem, both theoretically and empirically. We then show that a specific correction to the data distribution can mitigate this issue. Based on these observations, we propose a new algorithm, DisCor, which computes an approximation to this optimal distribution and uses it to re-weight the transitions used for training, resulting in substantial improvements in a range of challenging RL settings, such as multi-task learning and learning from noisy reward signals. Blog post presenting a summary of this work is available at: https://bair.berkeley.edu/blog/2020/03/16/discor/.
Pre-print
References in corpus (14)
- Rainbow: Combining Improvements in Deep Reinforcement Learning
- Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
- Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
- Behavior Regularized Offline Reinforcement Learning
- Dopamine: A Research Framework for Deep Reinforcement Learning
- Gradient Surgery for Multi-Task Learning
- Combating Label Noise in Deep Learning Using Abstention
- Provably Efficient Maximum Entropy Exploration
- Towards Characterizing Divergence in Deep Q-Learning
- Ray Interference: a Source of Plateaus in Deep Reinforcement Learning
- Approximate Policy Iteration Schemes: A Comparison
- Provably Efficient -learning with Function Approximation via Distribution Shift Error Checking Oracle
- Diagnosing Bottlenecks in Deep Q-learning Algorithms
- Tight Performance Bounds for Approximate Modified Policy Iteration with Non-Stationary Policies
Cited by in corpus (18)
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Conservative Q-Learning for Offline Reinforcement Learning
- MOPO: Model-based Offline Policy Optimization
- What are the Statistical Limits of Offline RL with Linear Function Approximation?
- Counterfactual Data Augmentation using Locally Factored Dynamics
- An Exponential Lower Bound for Linearly-Realizable MDPs with Constant Suboptimality Gap
- Shaping Rewards for Reinforcement Learning with Imperfect Demonstrations using Generative Models
- Implicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement Learning
- Qgraph-bounded Q-learning: Stabilizing Model-Free Off-Policy Deep Reinforcement Learning
- Robust Multi-Modal Policies for Industrial Assembly via Reinforcement Learning and Demonstrations: A Large-Scale Study
- GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning
- Conservative Data Sharing for Multi-Task Offline Reinforcement Learning
- Towards Automatic Actor-Critic Solutions to Continuous Control
- An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning
- Off-policy Reinforcement Learning with Optimistic Exploration and Distribution Correction
- Learning Pessimism for Robust and Efficient Off-Policy Reinforcement Learning
- A Closer Look at Advantage-Filtered Behavioral Cloning in High-Noise Datasets
- Generating GPU Compiler Heuristics using Reinforcement Learning