Hyper-Sphere Quantization: Communication-Efficient SGD for Federated Learning
arXiv:1911.04655
Abstract
The high cost of communicating gradients is a major bottleneck for federated learning, as the bandwidth of the participating user devices is limited. Existing gradient compression algorithms are mainly designed for data centers with high-speed network and achieve per-iteration communication cost at best, where is the size of the model. We propose hyper-sphere quantization (HSQ), a general framework that can be configured to achieve a continuum of trade-offs between communication efficiency and gradient accuracy. In particular, at the high compression ratio end, HSQ provides a low per-iteration communication cost of , which is favorable for federated learning. We prove the convergence of HSQ theoretically and show by experiments that HSQ significantly reduces the communication cost of model training without hurting convergence accuracy.
References in corpus (5)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Error Feedback Fixes SignSGD and other Gradient Compression Schemes
- AdaComp : Adaptive Residual Gradient Compression for Data-Parallel Distributed Training
- Distributed Learning with Sublinear Communication