1 paper
Xinran Gu, Kaifeng Lyu, Longbo Huang +1
Local SGD is a communication-efficient variant of SGD for large-scale training, where multiple GPUs perform SGD independently and average the model parameters periodically. It has…