activity
20172022
most citedConvergence Analysis of Proximal Gradient with Momentum for Nonconvex Optimization

36 citations · 193 across the 18 of their papers we have counts for

collaborators

38 papers

cs.LG2022

Finite-Time Error Bounds for Greedy-GQ

Yue Wang, Yi Zhou, Shaofeng Zou

Greedy-GQ with linear function approximation, originally proposed in \cite{maei2010toward}, is a value-based off-policy algorithm for optimal control in reinforcement learning, and…

math.OC2022★ 3 cited

On Unbalanced Optimal Transport: Gradient Methods, Sparsity and Approximation Error

Quang Minh Nguyen, Hoang H. Nguyen, Yi Zhou +1

We study the Unbalanced Optimal Transport (UOT) between two measures of possibly different masses with at most components, where the marginal constraints of standard Optimal Tr…

cs.LG2021★ 4 cited

Sample and Communication-Efficient Decentralized Actor-Critic Algorithms with Finite-Time Analysis

Ziyi Chen, Yi Zhou, Rongrong Chen +1

Actor-critic (AC) algorithms have been widely adopted in decentralized multi-agent systems to learn the optimal joint control policy. However, existing decentralized AC algorithms…

cs.LG2021

Non-Asymptotic Analysis for Two Time-scale TDC with General Smooth Function Approximation

Yue Wang, Shaofeng Zou, Yi Zhou

Temporal-difference learning with gradient correction (TDC) is a two time-scale algorithm for policy evaluation in reinforcement learning. This algorithm was initially proposed wit…

cs.LG2021★ 7 cited

Greedy-GQ with Variance Reduction: Finite-time Analysis and Improved Complexity

Shaocong Ma, Ziyi Chen, Yi Zhou +1

Greedy-GQ is a value-based reinforcement learning (RL) algorithm for optimal control. Recently, the finite-time analysis of Greedy-GQ has been developed under linear function appro…

math.OC2021★ 8 cited

Proximal Gradient Descent-Ascent: Variable Convergence under KŁ Geometry

Ziyi Chen, Yi Zhou, Tengyu Xu +1

The gradient descent-ascent (GDA) algorithm has been widely applied to solve minimax optimization problems. In order to achieve convergent policy parameters for minimax optimizatio…