2 papers
cs.LG2023
Evaluation and Optimization of Gradient Compression for Distributed Deep Learning
Lin Zhang, Longteng Zhang, Shaohuai Shi +2
To accelerate distributed training, many gradient compression methods have been proposed to alleviate the communication bottleneck in synchronous stochastic gradient descent (S-SGD…
cs.DC2021
Accelerating Distributed K-FAC with Smart Parallelism of Computing and Communication Tasks
Shaohuai Shi, Lin Zhang, Bo Li
Distributed training with synchronous stochastic gradient descent (SGD) on GPU clusters has been widely used to accelerate the training process of deep models. However, SGD only ut…