1 paper
Xin Zhou, Ling Chen, Houming Wu
As the size of models and datasets grows, it has become increasingly common to train models in parallel. However, existing distributed stochastic gradient descent (SGD) algorithms…