12 citations · 27 across the 5 of their papers we have counts for
5 papers · 1 filter
Revisiting Optimal Convergence Rate for Smooth and Non-convex Stochastic Decentralized Optimization
Kun Yuan, Xinmeng Huang, Yiming Chen +3
Decentralized optimization is effective to save communication in large-scale machine learning. Although numerous algorithms have been proposed with theoretical guarantees and empir…
Exponential Graph is Provably Efficient for Decentralized Deep Training
Bicheng Ying, Kun Yuan, Yiming Chen +3
Decentralized SGD is an emerging training method for deep learning known for its much less (thus faster) communication per iteration, which relaxes the averaging step in parallel S…
Accelerating Gossip SGD with Periodic Global Averaging
Yiming Chen, Kun Yuan, Yingya Zhang +3
Communication overhead hinders the scalability of large-scale distributed training. Gossip SGD, where each node averages only with its neighbors, is more communication-efficient th…
DecentLaM: Decentralized Momentum SGD for Large-batch Deep Training
Kun Yuan, Yiming Chen, Xinmeng Huang +4
The scale of deep learning nowadays calls for efficient distributed training algorithms. Decentralized momentum SGD (DmSGD), in which each node averages only with its neighbors, is…
Learned Gradient Compression for Distributed Deep Learning
Lusine Abrahamyan, Yiming Chen, Giannis Bekoulis +1
Training deep neural networks on large datasets containing high-dimensional data requires a large amount of computation. A solution to this problem is data-parallel distributed tra…