activity
20092022
most citedAn Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient Methods

30 citations · 106 across the 23 of their papers we have counts for

collaborators
Showing cs.LGShow all

10 papers · 1 filter

cs.LG202230 cited

An Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient Methods

Yanli Liu, Kaiqing Zhang, Tamer Başar +1

In this paper, we revisit and improve the convergence of policy gradient (PG), natural PG (NPG) methods, and their variance-reduced variants, under general smooth policy parametriz…

cs.LG2020

Fully Asynchronous Policy Evaluation in Distributed Reinforcement Learning over Networks

Xingyu Sha, Jiaqi Zhang, Keyou You +2

This paper proposes a \emph{fully asynchronous} scheme for the policy evaluation problem of distributed reinforcement learning (DisRL) over directed peer-to-peer networks. Without…

cs.LG201919 cited

Decentralized Multi-Agent Reinforcement Learning with Networked Agents: Recent Advances

Kaiqing Zhang, Zhuoran Yang, Tamer Başar

Multi-agent reinforcement learning (MARL) has long been a significant and everlasting research topic in both machine learning and control. With the recent development of (single-ag…

cs.LG2019

Online Planning for Decentralized Stochastic Control with Partial History Sharing

Kaiqing Zhang, Erik Miehling, Tamer Başar

In decentralized stochastic control, standard approaches for sequential decision-making, e.g. dynamic programming, quickly become intractable due to the need to maintain a complex…

cs.LG20192 cited

A Communication-Efficient Multi-Agent Actor-Critic Algorithm for Distributed Reinforcement Learning

Yixuan Lin, Kaiqing Zhang, Zhuoran Yang +4

This paper considers a distributed reinforcement learning problem in which a network of multiple agents aim to cooperatively maximize the globally averaged return through communica…

cs.LG2019

A Multi-Agent Off-Policy Actor-Critic Algorithm for Distributed Reinforcement Learning

Wesley Suttle, Zhuoran Yang, Kaiqing Zhang +3

This paper extends off-policy reinforcement learning to the multi-agent case in which a set of networked agents communicating with their neighbors according to a time-varying graph…