Q() with Off-Policy Corrections
arXiv:1602.04951
Abstract
We propose and analyze an alternate approach to off-policy multi-step temporal difference learning, in which off-policy returns are corrected with the current Q-function in terms of rewards, rather than with the target policy in terms of transition probabilities. We prove that such approximate corrections are sufficient for off-policy convergence both in policy evaluation and control, provided certain conditions. These conditions relate the distance between the target and behavior policies, the eligibility trace parameter and the discount factor, and formalize an underlying tradeoff in off-policy TD(). We illustrate this theoretical relationship empirically on a continuous-state control task.
Cited by in corpus (14)
- Safe and Efficient Off-Policy Reinforcement Learning
- Qualitative Measurements of Policy Discrepancy for Return-Based Deep Q-Network
- Sample-Efficient Deep Reinforcement Learning via Episodic Backward Update
- Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement Learning
- A Lyapunov Theory for Finite-Sample Guarantees of Asynchronous Q-Learning and TD-Learning Variants
- Assumed Density Filtering Q-learning
- Taylor Expansion Policy Optimization
- Finite-Sample Analysis of Off-Policy TD-Learning via Generalized Bellman Operators
- A survey of benchmarking frameworks for reinforcement learning
- Supervised Off-Policy Ranking
- Semi-On-Policy Training for Sample Efficient Multi-Agent Policy Gradients
- Learning Memory-Dependent Continuous Control from Demonstrations
- Context-aware Active Multi-Step Reinforcement Learning
- Gradient Q: A Unified Algorithm with Function Approximation for Reinforcement Learning