Safe and Efficient Off-Policy Reinforcement Learning
arXiv:1606.02647
Abstract
In this work, we take a fresh look at some old and new algorithms for off-policy, return-based reinforcement learning. Expressing these in a common form, we derive a novel algorithm, Retrace(), with three desired properties: (1) it has low variance; (2) it safely uses samples collected from any behaviour policy, whatever its degree of "off-policyness"; and (3) it is efficient as it makes the best use of samples collected from near on-policy behaviour policies. We analyze the contractive nature of the related operator under both off-policy policy evaluation and control settings and derive online sample-based algorithms. We believe this is the first return-based off-policy control algorithm converging a.s. to without the GLIE assumption (Greedy in the Limit with Infinite Exploration). As a corollary, we prove the convergence of Watkins' Q(), which was an open problem since 1989. We illustrate the benefits of Retrace() on a standard suite of Atari 2600 games.
References in corpus (3)
Cited by in corpus (19)
- Sample Efficient Actor-Critic with Experience Replay
- Neural Episodic Control
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning
- TD-Regularized Actor-Critic Methods
- Bridging adaptive management and reinforcement learning for more robust decisions
- Classification with Costly Features as a Sequential Decision-Making Problem
- Optimistic Reinforcement Learning by Forward Kullback-Leibler Divergence Optimization
- Experience Replay with Likelihood-free Importance Weights
- Estimation Error Correction in Deep Reinforcement Learning for Deterministic Actor-Critic Methods
- Optimizing Sequential Medical Treatments with Auto-Encoding Heuristic Search in POMDPs
- Interpolated Policy Gradient: Merging On-Policy and Off-Policy Gradient Estimation for Deep Reinforcement Learning
- A review of motion planning algorithms for intelligent robotics
- Towards Off-Policy Reinforcement Learning for Ranking Policies with Human Feedback
- Optimizing Medical Treatment for Sepsis in Intensive Care: from Reinforcement Learning to Pre-Trial Evaluation
- Investigating Recurrence and Eligibility Traces in Deep Q-Networks
- Diluted Near-Optimal Expert Demonstrations for Guiding Dialogue Stochastic Policy Optimisation
- Classification with Costly Features in Hierarchical Deep Sets
- Reinforcement Explanation Learning
- Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning