activity
20182021
most citedDouZero: Mastering DouDizhu with Self-Play Deep Reinforcement Learning

21 citations · 29 across the 3 of their papers we have counts for

collaborators

9 papers

cs.AI202121 cited

DouZero: Mastering DouDizhu with Self-Play Deep Reinforcement Learning

Daochen Zha, Jingru Xie, Wenye Ma +4

Games are abstractions of the real world, where artificial agents learn to compete and cooperate with other agents. While significant achievements have been made in various perfect…

cs.LG2021

1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence Speed

Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan +6

Scalable training of large models (like BERT and GPT-3) requires careful optimization rooted in model design, architecture, and system capabilities. From a system standpoint, commu…

cs.DC20204 cited

APMSqueeze: A Communication Efficient Adam-Preconditioned Momentum SGD Algorithm

Hanlin Tang, Shaoduo Gan, Samyam Rajbhandari +4

Adam is the important optimization algorithm to guarantee efficiency and accuracy for training many important tasks such as BERT and ImageNet. However, Adam is generally not compat…

stat.ML2020

Stochastic Recursive Momentum for Policy Gradient Methods

Huizhuo Yuan, Xiangru Lian, Ji Liu +1

In this paper, we propose a novel algorithm named STOchastic Recursive Momentum for Policy Gradient (STORM-PG), which operates a SARAH-type stochastic recursive variance-reduced po…

stat.ML20204 cited

Stochastic Recursive Variance Reduction for Efficient Smooth Non-Convex Compositional Optimization

Huizhuo Yuan, Xiangru Lian, Ji Liu

Stochastic compositional optimization arises in many important machine learning tasks such as value function evaluation in reinforcement learning and portfolio management. The obje…

cs.DC2019

: Decentralization Meets Error-Compensated Compression

Hanlin Tang, Xiangru Lian, Shuang Qiu +4

Communication is a key bottleneck in distributed training. Recently, an \emph{error-compensated} compression technology was particularly designed for the \emph{centralized} learnin…