most citedCredit Assignment with Meta-Policy Gradient for Multi-Agent Reinforcement Learning

5 citations · 12 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG20231 cited

Supported Trust Region Optimization for Offline Reinforcement Learning

Yixiu Mao, Hongchang Zhang, Chen Chen +2

Offline reinforcement learning suffers from the out-of-distribution issue and extrapolation error. Most policy constraint methods regularize the density of the trained policy towar…

cs.AI20232 cited

Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement Learning

Jianzhun Shao, Yun Qu, Chen Chen +2

Offline multi-agent reinforcement learning is challenging due to the coupling effect of both distribution shift issue common in offline setting and the high dimension issue common…

cs.LG20213 cited

Wasserstein Unsupervised Reinforcement Learning

Shuncheng He, Yuhang Jiang, Hongchang Zhang +2

Unsupervised reinforcement learning aims to train agents to learn a handful of policies or skills in environments without external reward. These pre-trained policies can accelerate…

cs.LG20211 cited

Reducing Conservativeness Oriented Offline Reinforcement Learning

Hongchang Zhang, Jianzhun Shao, Yuhang Jiang +2

In offline reinforcement learning, a policy learns to maximize cumulative rewards with a fixed collection of data. Towards conservative strategy, current methods choose to regulari…

cs.LG20215 cited

Credit Assignment with Meta-Policy Gradient for Multi-Agent Reinforcement Learning

Jianzhun Shao, Hongchang Zhang, Yuhang Jiang +2

Reward decomposition is a critical problem in centralized training with decentralized execution~(CTDE) paradigm for multi-agent reinforcement learning. To take full advantage of gl…