On Communication Compression for Distributed Optimization on Heterogeneous Data
arXiv:2009.02388
Abstract
Lossy gradient compression, with either unbiased or biased compressors, has become a key tool to avoid the communication bottleneck in centrally coordinated distributed training of machine learning models. We analyze the performance of two standard and general types of methods: (i) distributed quantized SGD (D-QSGD) with arbitrary unbiased quantizers and (ii) distributed SGD with error-feedback and biased compressors (D-EF-SGD) in the heterogeneous (non-iid) data setting. Our results indicate that D-EF-SGD is much less affected than D-QSGD by non-iid data, but both methods can suffer a slowdown if data-skewness is high. We further study two alternatives that are not (or much less) affected by heterogenous data distributions: first, a recently proposed method that is effective on strongly convex problems, and secondly, we point out a more general approach that is applicable to linear compressors only but effective in all considered scenarios.
References in corpus (9)
- Decentralized Stochastic Optimization and Gossip Algorithms with Compressed Communication
- A Unified Theory of Decentralized SGD with Changing Topology and Local Updates
- Error Feedback Fixes SignSGD and other Gradient Compression Schemes
- Stochastic Distributed Learning with Gradient Quantization and Variance Reduction
- FetchSGD: Communication-Efficient Federated Learning with Sketching
- On Biased Compression for Distributed Learning
- Linearly Converging Error Compensated SGD
- PowerGossip: Practical Low-Rank Communication Compression in Decentralized Deep Learning
- Trajectory Normalized Gradients for Distributed Optimization