42 citations · 112 across the 8 of their papers we have counts for
15 papers
Faster Algorithm and Sharper Analysis for Constrained Markov Decision Process
Tianjiao Li, Ziwei Guan, Shaofeng Zou +3
The problem of constrained Markov decision process (CMDP) is investigated, where an agent aims to maximize the expected accumulated discounted reward subject to multiple constraint…
A Unified Off-Policy Evaluation Approach for General Value Function
Tengyu Xu, Zhuoran Yang, Zhaoran Wang +1
General Value Function (GVF) is a powerful tool to represent both the {\em predictive} and {\em retrospective} knowledge in reinforcement learning (RL). In practice, often multiple…
Proximal Gradient Descent-Ascent: Variable Convergence under KŁ Geometry
Ziyi Chen, Yi Zhou, Tengyu Xu +1
The gradient descent-ascent (GDA) algorithm has been widely applied to solve minimax optimization problems. In order to achieve convergent policy parameters for minimax optimizatio…
Doubly Robust Off-Policy Actor-Critic: Convergence and Optimality
Tengyu Xu, Zhuoran Yang, Zhaoran Wang +1
Designing off-policy reinforcement learning algorithms is typically a very challenging task, because a desirable iteration update often involves an expectation over an on-policy di…
Sample Complexity Bounds for Two Timescale Value-based Reinforcement Learning Algorithms
Tengyu Xu, Yingbin Liang
Two timescale stochastic approximation (SA) has been widely used in value-based reinforcement learning algorithms. In the policy evaluation setting, it can model the linear and non…
CRPO: A New Approach for Safe Reinforcement Learning with Convergence Guarantee
Tengyu Xu, Yingbin Liang, Guanghui Lan
In safe reinforcement learning (SRL) problems, an agent explores the environment to maximize an expected total reward and meanwhile avoids violation of certain constraints on a num…