A Kernel Loss for Solving the Bellman Equation
arXiv:1905.10506
Abstract
Value function learning plays a central role in many state-of-the-art reinforcement-learning algorithms. Many popular algorithms like Q-learning do not optimize any objective function, but are fixed-point iterations of some variant of Bellman operator that is not necessarily a contraction. As a result, they may easily lose convergence guarantees, as can be observed in practice. In this paper, we propose a novel loss function, which can be optimized using standard gradient-based methods without risking divergence. The key advantage is that its gradient can be easily approximated using sampled transitions, avoiding the need for double samples required by prior algorithms like residual gradient. Our approach may be combined with general function classes such as neural networks, on either on- or off-policy data, and is shown to work reliably and effectively in several benchmarks.
17 pages, 5 figures, NeurIPS 2019
Cited by in corpus (21)
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- Minimax Weight and Q-Function Learning for Off-Policy Evaluation
- Causal Inference Under Unmeasured Confounding With Negative Controls: A Minimax Learning Approach
- Q* Approximation Schemes for Batch Reinforcement Learning: A Theoretical Comparison
- Minimax Value Interval for Off-Policy Evaluation and Policy Optimization
- Finite Sample Analysis of Minimax Offline Reinforcement Learning: Completeness, Fast Rates and First-Order Efficiency
- Doubly Robust Bias Reduction in Infinite Horizon Off-Policy Estimation
- Instrumental Variable Regression via Kernel Maximum Moment Loss
- Kalman meets Bellman: Improving Policy Evaluation through Value Tracking
- Deep Residual Reinforcement Learning
- Accountable Off-Policy Evaluation With Kernel Bellman Statistics
- Convergent and Efficient Deep Q Network Algorithm
- Optimal policy evaluation using kernel-based temporal difference methods
- Convex Q-Learning, Part 1: Deterministic Optimal Control
- A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
- Towards a practical measure of interference for reinforcement learning
- Non-asymptotic Confidence Intervals of Off-policy Evaluation: Primal and Dual Bounds
- Logistic Q-Learning
- A maximum-entropy approach to off-policy evaluation in average-reward MDPs
- Off-Policy Interval Estimation with Lipschitz Value Iteration
- Minimax Model Learning