activity
20182021
most citedTwo Time-scale Off-Policy TD Learning: Non-asymptotic Analysis over Markovian Samples

42 citations · 112 across the 8 of their papers we have counts for

collaborators

15 papers

math.OC20213 cited

Faster Algorithm and Sharper Analysis for Constrained Markov Decision Process

Tianjiao Li, Ziwei Guan, Shaofeng Zou +3

The problem of constrained Markov decision process (CMDP) is investigated, where an agent aims to maximize the expected accumulated discounted reward subject to multiple constraint…

cs.LG20211 cited

A Unified Off-Policy Evaluation Approach for General Value Function

Tengyu Xu, Zhuoran Yang, Zhaoran Wang +1

General Value Function (GVF) is a powerful tool to represent both the {\em predictive} and {\em retrospective} knowledge in reinforcement learning (RL). In practice, often multiple…

math.OC20218 cited

Proximal Gradient Descent-Ascent: Variable Convergence under KŁ Geometry

Ziyi Chen, Yi Zhou, Tengyu Xu +1

The gradient descent-ascent (GDA) algorithm has been widely applied to solve minimax optimization problems. In order to achieve convergent policy parameters for minimax optimizatio…

cs.LG2021

Doubly Robust Off-Policy Actor-Critic: Convergence and Optimality

Tengyu Xu, Zhuoran Yang, Zhaoran Wang +1

Designing off-policy reinforcement learning algorithms is typically a very challenging task, because a desirable iteration update often involves an expectation over an on-policy di…

cs.LG20205 cited

Sample Complexity Bounds for Two Timescale Value-based Reinforcement Learning Algorithms

Tengyu Xu, Yingbin Liang

Two timescale stochastic approximation (SA) has been widely used in value-based reinforcement learning algorithms. In the policy evaluation setting, it can model the linear and non…

cs.LG2020

CRPO: A New Approach for Safe Reinforcement Learning with Convergence Guarantee

Tengyu Xu, Yingbin Liang, Guanghui Lan

In safe reinforcement learning (SRL) problems, an agent explores the environment to maximize an expected total reward and meanwhile avoids violation of certain constraints on a num…