activity
20212024
most citedProvably Efficient Convergence of Primal-Dual Actor-Critic with Nonlinear Function Approximation

1 citations · 2 across the 5 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2023

DPMAC: Differentially Private Communication for Cooperative Multi-Agent Reinforcement Learning

Canzhe Zhao, Yanjie Ze, Jing Dong +2

Communication lays the foundation for cooperation in human society and in multi-agent reinforcement learning (MARL). Humans also desire to maintain their privacy when communicating…

cs.LG2023

Semantically Aligned Task Decomposition in Multi-Agent Reinforcement Learning

Wenhao Li, Dan Qiao, Baoxiang Wang +3

The difficulty of appropriately assigning credit is particularly heightened in cooperative MARL with sparse reward, due to the concurrent time and structural scales involved. Autom…

cs.LG20221 cited

Online Policy Optimization for Robust MDP

Jing Dong, Jingwei Li, Baoxiang Wang +1

Reinforcement learning (RL) has exceeded human performance in many synthetic settings such as video games and Go. However, real-world deployment of end-to-end RL models is less com…

cs.LG20221 cited

Provably Efficient Convergence of Primal-Dual Actor-Critic with Nonlinear Function Approximation

Jing Dong, Li Shen, Yinggan Xu +1

We study the convergence of the actor-critic algorithm with nonlinear function approximation under a nonconvex-nonconcave primal-dual formulation. Stochastic gradient descent ascen…

cs.LG2022

Differentially Private Temporal Difference Learning with Stochastic Nonconvex-Strongly-Concave Optimization

Canzhe Zhao, Yanjie Ze, Jing Dong +2

Temporal difference (TD) learning is a widely used method to evaluate policies in reinforcement learning. While many TD learning methods have been developed in recent years, little…

cs.LG2021

Incentivizing an Unknown Crowd

Jing Dong, Shuai Li, Baoxiang Wang

Motivated by the common strategic activities in crowdsourcing labeling, we study the problem of sequential eliciting information without verification (EIWV) for workers with a hete…