101 citations · 142 across the 7 of their papers we have counts for
7 papers
Revisiting Optimal Convergence Rate for Smooth and Non-convex Stochastic Decentralized Optimization
Kun Yuan, Xinmeng Huang, Yiming Chen +3
Decentralized optimization is effective to save communication in large-scale machine learning. Although numerous algorithms have been proposed with theoretical guarantees and empir…
Accelerating Gossip SGD with Periodic Global Averaging
Yiming Chen, Kun Yuan, Yingya Zhang +3
Communication overhead hinders the scalability of large-scale distributed training. Gossip SGD, where each node averages only with its neighbors, is more communication-efficient th…
DecentLaM: Decentralized Momentum SGD for Large-batch Deep Training
Kun Yuan, Yiming Chen, Xinmeng Huang +4
The scale of deep learning nowadays calls for efficient distributed training algorithms. Decentralized momentum SGD (DmSGD), in which each node averages only with its neighbors, is…
Large-Scale Training System for 100-Million Classification at Alibaba
Liuyihan Song, Pan Pan, Kang Zhao +5
In the last decades, extreme classification has become an essential topic for deep learning. It has achieved great success in many areas, especially in computer vision and natural…
Distribution Adaptive INT8 Quantization for Training CNNs
Kang Zhao, Sida Huang, Pan Pan +4
Researches have demonstrated that low bit-width (e.g., INT8) quantization can be employed to accelerate the inference process. It makes the gradient quantization very promising sin…
Visual Search at Alibaba
Yanhao Zhang, Pan Pan, Yun Zheng +4
This paper introduces the large scale visual search algorithm and system infrastructure at Alibaba. The following challenges are discussed under the E-commercial circumstance at Al…