101 citations · 230 across the 18 of their papers we have counts for
5 papers · 1 filter
Exponential Graph is Provably Efficient for Decentralized Deep Training
Bicheng Ying, Kun Yuan, Yiming Chen +3
Decentralized SGD is an emerging training method for deep learning known for its much less (thus faster) communication per iteration, which relaxes the averaging step in parallel S…
Accelerating Gossip SGD with Periodic Global Averaging
Yiming Chen, Kun Yuan, Yingya Zhang +3
Communication overhead hinders the scalability of large-scale distributed training. Gossip SGD, where each node averages only with its neighbors, is more communication-efficient th…
OR-Net: Pointwise Relational Inference for Data Completion under Partial Observation
Qianyu Feng, Linchao Zhu, Bang Zhang +2
Contemporary data-driven methods are typically fed with full supervision on large-scale datasets which limits their applicability. However, in the actual systems with limitations s…
DecentLaM: Decentralized Momentum SGD for Large-batch Deep Training
Kun Yuan, Yiming Chen, Xinmeng Huang +4
The scale of deep learning nowadays calls for efficient distributed training algorithms. Decentralized momentum SGD (DmSGD), in which each node averages only with its neighbors, is…
Large-Scale Training System for 100-Million Classification at Alibaba
Liuyihan Song, Pan Pan, Kang Zhao +5
In the last decades, extreme classification has become an essential topic for deep learning. It has achieved great success in many areas, especially in computer vision and natural…