On the Convergence Rate of Off-Policy Policy Optimization Methods with Density-Ratio Correction
arXiv:2106.00993
Abstract
In this paper, we study the convergence properties of off-policy policy improvement algorithms with state-action density ratio correction under function approximation setting, where the objective function is formulated as a max-max-min optimization problem. We characterize the bias of the learning objective and present two strategies with finite-time convergence guarantees. In our first strategy, we present algorithm P-SREDA with convergence rate , whose dependency on is optimal. In our second strategy, we propose a new off-policy actor-critic style algorithm named O-SPIM. We prove that O-SPIM converges to a stationary point with total complexity , which matches the convergence rate of some recent actor-critic algorithms in the on-policy setting.
48 Pages; AISTATS 2022
References in corpus (9)
- Finite-Sample Analysis of Proximal Gradient TD Algorithms
- AlgaeDICE: Policy Gradient from Arbitrary Experience
- GenDICE: Generalized Offline Estimation of Stationary Values
- Information-Theoretic Considerations in Batch Reinforcement Learning
- Off-Policy Policy Gradient with State Distribution Correction
- Non-asymptotic Convergence Analysis of Two Time-scale (Natural) Actor-Critic Algorithms
- Off-Policy Evaluation via the Regularized Lagrangian
- Variance-Reduced Off-Policy Memory-Efficient Policy Search
- Momentum-Based Policy Gradient Methods