Communication Compression for Decentralized Training
arXiv:1803.06443
Abstract
Optimizing distributed learning systems is an art of balancing between computation and communication. There have been two lines of research that try to deal with slower networks: {\em communication compression} for low bandwidth networks, and {\em decentralization} for high latency networks. In this paper, We explore a natural question: {\em can the combination of both techniques lead to a system that is robust to both bandwidth and latency?} Although the system implication of such combination is trivial, the underlying theoretical principle and algorithm design is challenging: unlike centralized algorithms, simply compressing exchanged information, even in an unbiased stochastic way, within the decentralized network would accumulate the error and fail to converge. In this paper, we develop a framework of compressed, decentralized training and propose two different strategies, which we call {\em extrapolation compression} and {\em difference compression}. We analyze both algorithms and prove both converge at the rate of where is the number of workers and is the number of iterations, matching the convergence rate for full precision, centralized training. We validate our algorithms and find that our proposed algorithm outperforms the best of merely decentralized and merely quantized algorithm significantly for networks with {\em both} high latency and low bandwidth.
Cited by in corpus (70)
- Decentralized Federated Learning: Fundamentals, State of the Art, Frameworks, Trends, and Challenges
- FedML: A Research Library and Benchmark for Federated Machine Learning
- Group Knowledge Transfer: Federated Learning of Large CNNs at the Edge
- A Unified Theory of Decentralized SGD with Changing Topology and Local Updates
- Database Meets Deep Learning: Challenges and Opportunities
- Federated Learning for Internet of Things: A Federated Learning Framework for On-device Anomaly Data Detection
- Edge Intelligence: Paving the Last Mile of Artificial Intelligence with Edge Computing
- Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
- Stochastic Distributed Learning with Gradient Quantization and Variance Reduction
- FedSAE: A Novel Self-Adaptive Federated Learning Framework in Heterogeneous Systems
- On the Influence of Bias-Correction on Distributed Stochastic Optimization
- Communication Efficiency in Federated Learning: Achievements and Challenges
- On-Device Machine Learning: An Algorithms and Learning Theory Perspective
- DataLens: Scalable Privacy Preserving Training via Gradient Compression and Aggregation
- MATCHA: Speeding Up Decentralized SGD via Matching Decomposition Sampling
- Communication-Censored Linearized ADMM for Decentralized Consensus Optimization
- Quantization for decentralized learning under subspace constraints
- Quasi-Global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous Data
- Hyper-Sphere Quantization: Communication-Efficient SGD for Federated Learning
- Moniqua: Modulo Quantized Communication in Decentralized SGD
- Hierarchical Federated Learning through LAN-WAN Orchestration
- Decentralized Deep Learning with Arbitrary Communication Compression
- A Double Residual Compression Algorithm for Efficient Distributed Learning
- Communication-Efficient Local Decentralized SGD Methods
- Compressed Gradient Tracking for Decentralized Optimization Over General Directed Networks
- Federated Learning with Compression: Unified Analysis and Sharp Guarantees
- Communication-Efficient Edge AI: Algorithms and Systems
- Decentralized Stochastic Gradient Tracking for Non-convex Empirical Risk Minimization
- : Decentralization Meets Error-Compensated Compression
- GFL: A Decentralized Federated Learning Framework Based On Blockchain
- On the Utility of Gradient Compression in Distributed Training Systems
- SPARQ-SGD: Event-Triggered and Compressed Communication in Decentralized Stochastic Optimization
- Linear Convergent Decentralized Optimization with Compression
- To Talk or to Work: Flexible Communication Compression for Energy Efficient Federated Learning over Heterogeneous Mobile Edge Devices
- FedSKETCH: Communication-Efficient and Private Federated Learning via Sketching
- Communication-Efficient Decentralized Learning with Sparsification and Adaptive Peer Selection
- Straggler-Resilient Distributed Machine Learning with Dynamic Backup Workers
- On Communication Compression for Distributed Optimization on Heterogeneous Data
- Faster Non-Convex Federated Learning via Global and Local Momentum
- TEE-based decentralized recommender systems: The raw data sharing redemption
- Periodic Stochastic Gradient Descent with Momentum for Decentralized Training
- Distributed Optimization over Block-Cyclic Data
- Error Compensated Distributed SGD Can Be Accelerated
- PowerGossip: Practical Low-Rank Communication Compression in Decentralized Deep Learning
- Heterogeneity-Aware Asynchronous Decentralized Training
- Improved Convergence Analysis and SNR Control Strategies for Federated Learning in the Presence of Noise
- Byzantine-resilient Decentralized Stochastic Gradient Descent
- Get More for Less in Decentralized Learning Systems
- Innovation Compression for Communication-efficient Distributed Optimization with Linear Convergence
- Private and Communication-Efficient Edge Learning: A Sparse Differential Gaussian-Masking Distributed SGD Approach
- On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Optimization
- Efficient Ring-topology Decentralized Federated Learning with Deep Generative Models for Industrial Artificial Intelligent
- NUQSGD: Provably Communication-efficient Data-parallel SGD via Nonuniform Quantization
- A Linearly Convergent Algorithm for Decentralized Optimization: Sending Less Bits for Free!
- Decentralized Composite Optimization with Compression
- Adaptive Serverless Learning
- Decentralized Optimization On Time-Varying Directed Graphs Under Communication Constraints
- Step-Ahead Error Feedback for Distributed Training with Compressed Gradient
- Optimal Complexity in Decentralized Training
- Federated Submodel Optimization for Hot and Cold Data Features
- Error Compensated Loopless SVRG, Quartz, and SDCA for Distributed Optimization
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
- CADA: Communication-Adaptive Distributed Adam
- Finite-Time Consensus Learning for Decentralized Optimization with Nonlinear Gossiping
- CDC: Classification Driven Compression for Bandwidth Efficient Edge-Cloud Collaborative Deep Learning
- On the Convergence of Quantized Parallel Restarted SGD for Central Server Free Distributed Training
- Communication Efficient Generalized Tensor Factorization for Decentralized Healthcare Networks
- Communication-Efficient Network-Distributed Optimization with Differential-Coded Compressors
- Bristle: Decentralized Federated Learning in Byzantine, Non-i.i.d. Environments
- An Empirical Study on Compressed Decentralized Stochastic Gradient Algorithms with Overparameterized Models