36 citations · 193 across the 18 of their papers we have counts for
38 papers
Finite-Time Error Bounds for Greedy-GQ
Yue Wang, Yi Zhou, Shaofeng Zou
Greedy-GQ with linear function approximation, originally proposed in \cite{maei2010toward}, is a value-based off-policy algorithm for optimal control in reinforcement learning, and…
On Unbalanced Optimal Transport: Gradient Methods, Sparsity and Approximation Error
Quang Minh Nguyen, Hoang H. Nguyen, Yi Zhou +1
We study the Unbalanced Optimal Transport (UOT) between two measures of possibly different masses with at most components, where the marginal constraints of standard Optimal Tr…
Sample and Communication-Efficient Decentralized Actor-Critic Algorithms with Finite-Time Analysis
Ziyi Chen, Yi Zhou, Rongrong Chen +1
Actor-critic (AC) algorithms have been widely adopted in decentralized multi-agent systems to learn the optimal joint control policy. However, existing decentralized AC algorithms…
Non-Asymptotic Analysis for Two Time-scale TDC with General Smooth Function Approximation
Yue Wang, Shaofeng Zou, Yi Zhou
Temporal-difference learning with gradient correction (TDC) is a two time-scale algorithm for policy evaluation in reinforcement learning. This algorithm was initially proposed wit…
Greedy-GQ with Variance Reduction: Finite-time Analysis and Improved Complexity
Shaocong Ma, Ziyi Chen, Yi Zhou +1
Greedy-GQ is a value-based reinforcement learning (RL) algorithm for optimal control. Recently, the finite-time analysis of Greedy-GQ has been developed under linear function appro…
Proximal Gradient Descent-Ascent: Variable Convergence under KŁ Geometry
Ziyi Chen, Yi Zhou, Tengyu Xu +1
The gradient descent-ascent (GDA) algorithm has been widely applied to solve minimax optimization problems. In order to achieve convergent policy parameters for minimax optimizatio…