36 citations · 38 across the 12 of their papers we have counts for
1 paper · 1 filter
Ashok Vardhan Makkuva, Marco Bondaschi, Thijs Vogels +3
Data-parallel SGD is the de facto algorithm for distributed optimization, especially for large scale machine learning. Despite its merits, communication bottleneck is one of its pe…