90 citations · 90 across the 1 of their papers we have counts for
3 papers
cs.LG2020
STL-SGD: Speeding Up Local SGD with Stagewise Communication Period
Shuheng Shen, Yifei Cheng, Jingchang Liu +1
Distributed parallel stochastic gradient descent algorithms are workhorses for large scale machine learning tasks. Among them, local stochastic gradient descent (Local SGD) has att…
cs.LG2019★ 90 cited
Variance Reduced Local SGD with Lower Communication Complexity
Xianfeng Liang, Shuheng Shen, Jingchang Liu +3
To accelerate the training of machine learning models, distributed stochastic gradient descent (SGD) and its variants have been widely adopted, which apply multiple workers in para…
cs.LG2019
Faster Distributed Deep Net Training: Computation and Communication Decoupled Stochastic Gradient Descent
Shuheng Shen, Linli Xu, Jingchang Liu +2
With the increase in the amount of data and the expansion of model scale, distributed parallel training becomes an important and successful technique to address the optimization ch…