18 citations · 18 across the 1 of their papers we have counts for
2 papers
cs.LG2020
Is Local SGD Better than Minibatch SGD?
Blake Woodworth, Kumar Kshitij Patel, Sebastian U. Stich +5
We study local SGD (also known as parallel SGD and federated averaging), a natural and frequently used stochastic distributed optimization method. Its theoretical foundations are c…
cs.LG2019★ 18 cited
Communication trade-offs for synchronized distributed SGD with large step size
Kumar Kshitij Patel, Aymeric Dieuleveut
Synchronous mini-batch SGD is state-of-the-art for large-scale distributed machine learning. However, in practice, its convergence is bottlenecked by slow communication rounds betw…