Sparse Communication for Distributed Gradient Descent
arXiv:1704.05021 · doi:10.18653/v1/D17-1045
Abstract
We make distributed stochastic gradient descent faster by exchanging sparse updates instead of dense updates. Gradient updates are positively skewed as most updates are near zero, so we map the 99% smallest updates (by absolute value) to zero then exchange sparse matrices. This method can be combined with quantization to further improve the compression. We explore different configurations and apply them to neural machine translation and MNIST image classification tasks. Most configurations work on MNIST, whereas different configurations reduce convergence rate on the more complex translation task. Our experiments show that we can achieve up to 49% speed up on MNIST and 22% on NMT without damaging the final accuracy or BLEU.
EMNLP 2017
Cited by in corpus (39)
- Machine Learning at the Wireless Edge: Distributed Stochastic Gradient Descent Over-the-Air
- Vertical Federated Learning: Concepts, Advances and Challenges
- Multi-Armed Bandit Based Client Scheduling for Federated Learning
- UVeQFed: Universal Vector Quantization for Federated Learning
- Over-the-Air Federated Learning from Heterogeneous Data
- Pervasive AI for IoT applications: A Survey on Resource-efficient Distributed Artificial Intelligence
- A Survey on Approximate Edge AI for Energy Efficient Autonomous Driving Services
- Joint Optimization of Communications and Federated Learning Over the Air
- Communication optimization strategies for distributed deep neural network training: A survey
- Joint Privacy Enhancement and Quantization in Federated Learning
- Computation Scheduling for Distributed Machine Learning with Straggling Workers
- On-board Federated Learning for Satellite Clusters with Inter-Satellite Links
- Communication-Efficient Federated Learning with Binary Neural Networks
- Privacy-preserving Incremental ADMM for Decentralized Consensus Optimization
- Towards Efficient Synchronous Federated Training: A Survey on System Optimization Strategies
- Graph Federated Learning for CIoT Devices in Smart Home Applications
- BEV-SGD: Best Effort Voting SGD for Analog Aggregation Based Federated Learning against Byzantine Attackers
- FedVQCS: Federated Learning via Vector Quantized Compressed Sensing
- Large-Scale Training System for 100-Million Classification at Alibaba
- Training Simplification and Model Simplification for Deep Learning: A Minimal Effort Back Propagation Method
- Learning Rate Optimization for Federated Learning Exploiting Over-the-air Computation
- Ternary Compression for Communication-Efficient Federated Learning
- JSDoop and TensorFlow.js: Volunteer Distributed Web Browser-Based Neural Network Training
- Boosting Distributed Machine Learning Training Through Loss-tolerant Transmission Protocol
- Lottery Hypothesis based Unsupervised Pre-training for Model Compression in Federated Learning
- ScionFL: Efficient and Robust Secure Quantized Aggregation
- Masked Random Noise for Communication Efficient Federated Learning
- Beyond Throughput and Compression Ratios: Towards High End-to-end Utility of Gradient Compression
- Get More for Less in Decentralized Learning Systems
- Data-Aware Gradient Compression for FL in Communication-Constrained Mobile Computing
- Compressing gradients by exploiting temporal correlation in momentum-SGD
- Sparse Training for Federated Learning with Regularized Error Correction
- VFLAIR-LLM: A Comprehensive Framework and Benchmark for Split Learning of LLMs
- Approximate Wireless Communication for Lossy Gradient Updates in IoT Federated Learning
- Regularized Top-: A Bayesian Framework for Gradient Sparsification
- Communication-Efficient Distributed Learning via Sparse and Adaptive Stochastic Gradient
- Revolutionizing Wireless Networks with Federated Learning: A Comprehensive Review
- Sparse Incremental Aggregation in Multi-Hop Federated Learning
- FedShift: Robust Federated Learning Aggregation Scheme in Resource Constrained Environment via Weight Shifting