QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
arXiv:1610.02132
Abstract
Parallel implementations of stochastic gradient descent (SGD) have received significant research attention, thanks to excellent scalability properties of this algorithm, and to its efficiency in the context of training deep neural networks. A fundamental barrier for parallelizing large-scale SGD is the fact that the cost of communicating the gradient updates between nodes can be very large. Consequently, lossy compression heuristics have been proposed, by which nodes only communicate quantized gradients. Although effective in practice, these heuristics do not always provably converge, and it is not clear whether they are optimal. In this paper, we propose Quantized SGD (QSGD), a family of compression schemes which allow the compression of gradient updates at each node, while guaranteeing convergence under standard assumptions. QSGD allows the user to trade off compression and convergence time: it can communicate a sublinear number of bits per iteration in the model dimension, and can achieve asymptotically optimal communication cost. We complement our theoretical results with empirical data, showing that QSGD can significantly reduce communication cost, while being competitive with standard uncompressed techniques on a variety of real tasks. In particular, experiments show that gradient quantization applied to training of deep neural networks for image classification and automated speech recognition can lead to significant reductions in communication cost, and end-to-end training time. For instance, on 16 GPUs, we are able to train a ResNet-152 network on ImageNet 1.8x faster to full accuracy. Of note, we show that there exist generic parameter settings under which all known network architectures preserve or slightly improve their full accuracy when using quantization.
Cited by in corpus (283)
- Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
- Machine Learning at the Wireless Edge: Distributed Stochastic Gradient Descent Over-the-Air
- Decentralized Federated Learning: Fundamentals, State of the Art, Frameworks, Trends, and Challenges
- Sustainable AI: Environmental Implications, Challenges and Opportunities
- FedPAQ: A Communication-Efficient Federated Learning Method with Periodic Averaging and Quantization
- UVeQFed: Universal Vector Quantization for Federated Learning
- Model compression via distillation and quantization
- Over-the-Air Federated Learning from Heterogeneous Data
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Gradient Sparsification for Communication-Efficient Distributed Optimization
- A Survey on Methods and Theories of Quantized Neural Networks
- HeteroFL: Computation and Communication Efficient Federated Learning for Heterogeneous Clients
- Group Knowledge Transfer: Federated Learning of Large CNNs at the Edge
- cpSGD: Communication-efficient and differentially-private distributed SGD
- Federated Learning: A Signal Processing Perspective
- A Field Guide to Federated Optimization
- A Unified Theory of Decentralized SGD with Changing Topology and Local Updates
- PowerSGD: Practical Low-Rank Gradient Compression for Distributed Optimization
- An Exact Quantized Decentralized Gradient Descent Algorithm
- High-Dimensional Stochastic Gradient Quantization for Communication-Efficient Edge Learning
- Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
- Distributed Learning with Compressed Gradient Differences
- Optimal Client Sampling for Federated Learning
- Mime: Mimicking Centralized Stochastic Algorithms in Federated Learning
- Cronus: Robust and Heterogeneous Collaborative Learning with Black-Box Knowledge Transfer
- Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
- Variance Reduced Local SGD with Lower Communication Complexity
- VAFL: a Method of Vertical Asynchronous Federated Learning
- Joint Optimization of Communications and Federated Learning Over the Air
- Fast Federated Learning by Balancing Communication Trade-Offs
- On Maintaining Linear Convergence of Distributed Learning and Optimization under Limited Communication
- LotteryFL: Personalized and Communication-Efficient Federated Learning with Lottery Ticket Hypothesis on Non-IID Datasets
- SlowMo: Improving Communication-Efficient Distributed SGD with Slow Momentum
- Understanding Top-k Sparsification in Distributed Deep Learning
- Communication optimization strategies for distributed deep neural network training: A survey
- Communication-Efficient Distributed Blockwise Momentum SGD with Error-Feedback
- Acceleration for Compressed Gradient Descent in Distributed and Federated Optimization
- Natural Compression for Distributed Deep Learning
- FetchSGD: Communication-Efficient Federated Learning with Sketching
- FedLab: A Flexible Federated Learning Framework
- Better Theory for SGD in the Nonconvex World
- Bi-GCN: Binary Graph Convolutional Network
- Near-Optimal Sparse Allreduce for Distributed Deep Learning
- Communication Efficient Federated Learning over Multiple Access Channels
- DFTerNet: Towards 2-bit Dynamic Fusion Networks for Accurate Human Activity Recognition
- On-Device Machine Learning: An Algorithms and Learning Theory Perspective
- Oort: Efficient Federated Learning via Guided Participant Selection
- On Biased Compression for Distributed Learning
- DataLens: Scalable Privacy Preserving Training via Gradient Compression and Aggregation
- The OARF Benchmark Suite: Characterization and Implications for Federated Learning Systems
- Towards Communication-efficient and Attack-Resistant Federated Edge Learning for Industrial Internet of Things
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- RedSync : Reducing Synchronization Traffic for Distributed Deep Learning
- Distributed learning with compressed gradients
- Linear Convergence in Federated Learning: Tackling Client Heterogeneity and Sparse Gradients
- SemiFL: Semi-Supervised Federated Learning for Unlabeled Clients with Alternate Training
- A Federated Deep Learning Framework for Privacy Preservation and Communication Efficiency
- Quantization for decentralized learning under subspace constraints
- Hyper-Sphere Quantization: Communication-Efficient SGD for Federated Learning
- Breaking the Communication-Privacy-Accuracy Trilemma
- Edge-assisted Democratized Learning Towards Federated Analytics
- AnycostFL: Efficient On-Demand Federated Learning over Heterogeneous Edge Devices
- Decentralized Deep Learning with Arbitrary Communication Compression
- Overfitting for Fun and Profit: Instance-Adaptive Data Compression
- ElasticTrainer: Speeding Up On-Device Training with Runtime Elastic Tensor Selection
- A Double Residual Compression Algorithm for Efficient Distributed Learning
- QUOTIENT: Two-Party Secure Neural Network Training and Prediction
- Compressed Gradient Tracking for Decentralized Optimization Over General Directed Networks
- A Comprehensive Review and a Taxonomy of Edge Machine Learning: Requirements, Paradigms, and Techniques
- Recent theoretical advances in decentralized distributed convex optimization
- Towards Scalable Distributed Training of Deep Learning on Public Cloud Clusters
- Convergence of Distributed Stochastic Variance Reduced Methods without Sampling Extra Data
- FEDZIP: A Compression Framework for Communication-Efficient Federated Learning
- : Decentralization Meets Error-Compensated Compression
- A Survey of Coded Distributed Computing
- Federated Accelerated Stochastic Gradient Descent
- Local AdaAlter: Communication-Efficient Stochastic Gradient Descent with Adaptive Learning Rates
- Towards Unified INT8 Training for Convolutional Neural Network
- Accuracy-Efficiency Trade-Offs and Accountability in Distributed ML Systems
- A Better Alternative to Error Feedback for Communication-Efficient Distributed Learning
- Rethinking gradient sparsification as total error minimization
- Differentially Private Federated Learning for Resource-Constrained Internet of Things
- Check-N-Run: A Checkpointing System for Training Deep Learning Recommendation Models
- FedGreen: Federated Learning with Fine-Grained Gradient Compression for Green Mobile Edge Computing
- Robust Distributed Optimization With Randomly Corrupted Gradients
- BROADCAST: Reducing Both Stochastic and Compression Noise to Robustify Communication-Efficient Federated Learning
- Compressive Sensing Using Iterative Hard Thresholding with Low Precision Data Representation: Theory and Applications
- A Unified Theory of SGD: Variance Reduction, Sampling, Quantization and Coordinate Descent
- Bayesian Federated Learning over Wireless Networks
- Secure Aggregation with Heterogeneous Quantization in Federated Learning
- FedPara: Low-Rank Hadamard Product for Communication-Efficient Federated Learning
- On the Convergence of SGD with Biased Gradients
- Accordion: Adaptive Gradient Communication via Critical Learning Regime Identification
- Federated Learning for Energy Constrained IoT devices: A systematic mapping study
- Decentralized Learning Made Easy with DecentralizePy
- The ZipML Framework for Training Models with End-to-End Low Precision: The Cans, the Cannots, and a Little Bit of Deep Learning
- MARINA: Faster Non-Convex Distributed Learning with Compression
- Distributed Learning of Deep Neural Networks using Independent Subnet Training
- Layer-wise Adaptive Gradient Sparsification for Distributed Deep Learning with Convergence Guarantees
- Learning Rate Optimization for Federated Learning Exploiting Over-the-air Computation
- Secure Aggregation for Buffered Asynchronous Federated Learning
- A Carbon Tracking Model for Federated Learning: Impact of Quantization and Sparsification
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- Parallel Restarted SPIDER -- Communication Efficient Distributed Nonconvex Optimization with Optimal Computation Complexity
- Linear Convergent Decentralized Optimization with Compression
- Domain-specific Communication Optimization for Distributed DNN Training
- Local SGD With a Communication Overhead Depending Only on the Number of Workers
- Boosting Distributed Machine Learning Training Through Loss-tolerant Transmission Protocol
- To Talk or to Work: Flexible Communication Compression for Energy Efficient Federated Learning over Heterogeneous Mobile Edge Devices
- Local Stochastic Gradient Descent Ascent: Convergence Analysis and Communication Efficiency
- Gradient Descent with Compressed Iterates
- rTop-k: A Statistical Estimation Approach to Distributed SGD
- BlueFog: Make Decentralized Algorithms Practical for Optimization and Deep Learning
- Communication Efficient Distributed Learning with Censored, Quantized, and Generalized Group ADMM
- Scalable and Communication-efficient Decentralized Federated Edge Learning with Multi-blockchain Framework
- FedSKETCH: Communication-Efficient and Private Federated Learning via Sketching
- DynaComm: Accelerating Distributed CNN Training between Edges and Clouds through Dynamic Communication Scheduling
- Quantized Frank-Wolfe: Faster Optimization, Lower Communication, and Projection Free
- Shuffled Model of Federated Learning: Privacy, Communication and Accuracy Trade-offs
- Matrix Exponential Learning Schemes with Low Informational Exchange
- Distributed Second Order Methods with Fast Rates and Compressed Communication
- GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training
- One-Bit Over-the-Air Aggregation for Communication-Efficient Federated Edge Learning: Design and Convergence Analysis
- Communication-Efficient Distributed SGD with Compressed Sensing
- Efficient Neural Architecture Search via Proximal Iterations
- Faster Non-Convex Federated Learning via Global and Local Momentum
- Accelerating DNN Training in Wireless Federated Edge Learning Systems
- DRIVE: One-bit Distributed Mean Estimation
- On Communication Compression for Distributed Optimization on Heterogeneous Data
- Federated Learning over Wireless Device-to-Device Networks: Algorithms and Convergence Analysis
- Byzantine-Resilient Secure Federated Learning
- Robust and Communication-Efficient Collaborative Learning
- Per-Tensor Fixed-Point Quantization of the Back-Propagation Algorithm
- Periodic Stochastic Gradient Descent with Momentum for Decentralized Training
- Large-Scale Deep Learning Optimizations: A Comprehensive Survey
- Channel-driven Decentralized Bayesian Federated Learning for Trustworthy Decision Making in D2D Networks
- Wyner-Ziv Gradient Compression for Federated Learning
- Linear Regression with Distributed Learning: A Generalization Error Perspective
- EventGraD: Event-Triggered Communication in Parallel Machine Learning
- Error Compensated Distributed SGD Can Be Accelerated
- Bidirectional compression in heterogeneous settings for distributed or federated learning with partial participation: tight convergence guarantees
- PowerGossip: Practical Low-Rank Communication Compression in Decentralized Deep Learning
- Privacy For Free: Wireless Federated Learning Via Uncoded Transmission With Adaptive Power Control
- Efficient Halftoning via Deep Reinforcement Learning
- Accelerating Neural Network Training with Distributed Asynchronous and Selective Optimization (DASO)
- Elastic Consistency: A General Consistency Model for Distributed Stochastic Gradient Descent
- Wireless Data Acquisition for Edge Learning: Data-Importance Aware Retransmission
- Nested Distributed Gradient Methods with Adaptive Quantized Communication
- Fairness-aware Agnostic Federated Learning
- Accelerated Training for CNN Distributed Deep Learning through Automatic Resource-Aware Layer Placement
- COKE: Communication-Censored Decentralized Kernel Learning
- CANITA: Faster Rates for Distributed Convex Optimization with Communication Compression
- GraVAC: Adaptive Compression for Communication-Efficient Distributed DL Training
- EmbRace: Accelerating Sparse Communication for Distributed Training of NLP Neural Networks
- Gossiped and Quantized Online Multi-Kernel Learning
- Activations and Gradients Compression for Model-Parallel Training
- Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks
- On the Convergence of Decentralized Adaptive Gradient Methods
- Dynamic Scheduling for Over-the-Air Federated Edge Learning with Energy Constraints
- Nested Dithered Quantization for Communication Reduction in Distributed Training
- Distributed Newton Can Communicate Less and Resist Byzantine Workers
- Pufferfish: Communication-efficient Models At No Extra Cost
- Distributed Training of Graph Convolutional Networks using Subgraph Approximation
- A Unified Analysis of Variational Inequality Methods: Variance Reduction, Sampling, Quantization and Coordinate Descent
- Adaptive Precision Training (AdaPT): A dynamic fixed point quantized training approach for DNNs
- 99% of Distributed Optimization is a Waste of Time: The Issue and How to Fix it
- Get More for Less in Decentralized Learning Systems
- Data-Importance Aware User Scheduling for Communication-Efficient Edge Machine Learning
- Communication-Efficient Federated Learning with Compensated Overlap-FedAvg
- Leveraging Spatial and Temporal Correlations in Sparsified Mean Estimation
- Federated Learning in Adversarial Settings
- On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Optimization
- Innovation Compression for Communication-efficient Distributed Optimization with Linear Convergence
- Trends and Advancements in Deep Neural Network Communication
- Decentralized Composite Optimization with Compression
- MG-WFBP: Efficient Data Communication for Distributed Synchronous SGD Algorithms
- Adaptive Serverless Learning
- PFA: Privacy-preserving Federated Adaptation for Effective Model Personalization
- A Closer Look at Codistillation for Distributed Training
- APMSqueeze: A Communication Efficient Adam-Preconditioned Momentum SGD Algorithm
- Rate-Distortion Theoretic Bounds on Generalization Error for Distributed Learning
- Communication Efficient Federated Learning with Energy Awareness over Wireless Networks
- Distributed Inexact Successive Convex Approximation ADMM: Analysis-Part I
- Distributed Optimization for Over-Parameterized Learning
- Faster Distributed Deep Net Training: Computation and Communication Decoupled Stochastic Gradient Descent
- Optimal Gradient Quantization Condition for Communication-Efficient Distributed Training
- New Bounds For Distributed Mean Estimation and Variance Reduction
- A flexible framework for communication-efficient machine learning: from HPC to IoT
- Daydream: Accurately Estimating the Efficacy of Optimizations for DNN Training
- Improving the convergence of SGD through adaptive batch sizes
- Straggler-Agnostic and Communication-Efficient Distributed Primal-Dual Algorithm for High-Dimensional Data Mining
- 1-Bit Compressive Sensing for Efficient Federated Learning Over the Air
- Moshpit SGD: Communication-Efficient Decentralized Training on Heterogeneous Unreliable Devices
- Sign Bit is Enough: A Learning Synchronization Framework for Multi-hop All-reduce with Ultimate Compression
- QLSD: Quantised Langevin stochastic dynamics for Bayesian federated learning
- Federated Learning is Better with Non-Homomorphic Encryption
- SuperNeurons: FFT-based Gradient Sparsification in the Distributed Training of Deep Neural Networks
- Communication-Efficient Zeroth-Order Distributed Online Optimization: Algorithm, Theory, and Applications
- Accelerating Gossip SGD with Periodic Global Averaging
- Compressing gradients by exploiting temporal correlation in momentum-SGD
- Communication-Efficient Policy Gradient Methods for Distributed Reinforcement Learning
- Distributed Machine Learning for Wireless Communication Networks: Techniques, Architectures, and Applications
- SGD with Coordinate Sampling: Theory and Practice
- LAGC: Lazily Aggregated Gradient Coding for Straggler-Tolerant and Communication-Efficient Distributed Learning
- ErrorCompensatedX: error compensation for variance reduced algorithms
- Quantized Epoch-SGD for Communication-Efficient Distributed Learning
- Communication-Efficient Distributed SGD using Preamble-based Random Access
- Inclusive Data Representation in Federated Learning: A Novel Approach Integrating Textual and Visual Prompt
- Permutation Compressors for Provably Faster Distributed Nonconvex Optimization
- Compressed Distributed Gradient Descent: Communication-Efficient Consensus over Networks
- About some works of Boris Polyak on convergence of gradient methods and their development
- Communication-Efficient Distributed Optimization with Quantized Preconditioners
- Making Asynchronous Stochastic Gradient Descent Work for Transformers
- Sync-Switch: Hybrid Parameter Synchronization for Distributed Deep Learning
- Boosted and Differentially Private Ensembles of Decision Trees
- A Hybrid-Order Distributed SGD Method for Non-Convex Optimization to Balance Communication Overhead, Computational Complexity, and Convergence Rate
- QuPeL: Quantized Personalization with Applications to Federated Learning
- FlexPD: A Flexible Framework Of First-Order Primal-Dual Algorithms for Distributed Optimization
- Federated Learning over Wireless Networks: A Band-limited Coordinated Descent Approach
- A Survey on Large-scale Machine Learning
- Adaptive Periodic Averaging: A Practical Approach to Reducing Communication in Distributed Learning
- Trajectory Normalized Gradients for Distributed Optimization
- Differential Privacy Meets Federated Learning under Communication Constraints
- Neural Tangent Kernel Empowered Federated Learning
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
- FedDQ: Communication-Efficient Federated Learning with Descending Quantization
- Solving Multi-Arm Bandit Using a Few Bits of Communication
- Asynchronous Stochastic Optimization Robust to Arbitrary Delays
- Slashing Communication Traffic in Federated Learning by Transmitting Clustered Model Updates
- Finite-Time Consensus Learning for Decentralized Optimization with Nonlinear Gossiping
- Communication-Efficient Distributed Learning via Sparse and Adaptive Stochastic Gradient
- FedProf: Selective Federated Learning with Representation Profiling
- Local Methods with Adaptivity via Scaling
- Towards Tight Communication Lower Bounds for Distributed Optimisation
- Compressed Communication for Distributed Training: Adaptive Methods and System
- CD-SGD: Distributed Stochastic Gradient Descent with Compression and Delay Compensation
- CFedAvg: Achieving Efficient Communication and Fast Convergence in Non-IID Federated Learning
- Non-asymptotic moment bounds for random variables rounded to non-uniformly spaced sets
- CSER: Communication-efficient SGD with Error Reset
- Error Compensated Loopless SVRG, Quartz, and SDCA for Distributed Optimization
- On the Convergence of Quantized Parallel Restarted SGD for Central Server Free Distributed Training
- CADA: Communication-Adaptive Distributed Adam
- FedNS: Improving Federated Learning for collaborative image classification on mobile clients
- Critical Parameters for Scalable Distributed Learning with Large Batches and Asynchronous Updates
- CrossoverScheduler: Overlapping Multiple Distributed Training Applications in a Crossover Manner
- Sparsification as a Remedy for Staleness in Distributed Asynchronous SGD
- Information-constrained optimization: can adaptive processing of gradients help?
- Communication-Censored Distributed Stochastic Gradient Descent
- Communication Efficient Federated Learning with Adaptive Quantization
- OD-SGD: One-step Delay Stochastic Gradient Descent for Distributed Training
- DEED: A General Quantization Scheme for Communication Efficiency in Bits
- Decentralized Learning with Lazy and Approximate Dual Gradients
- WOR and 's: Sketches for -Sampling Without Replacement
- Coded Stochastic ADMM for Decentralized Consensus Optimization with Edge Computing
- MG-WFBP: Merging Gradients Wisely for Efficient Communication in Distributed Deep Learning
- Optimal Compression of Locally Differentially Private Mechanisms
- Design and Analysis of Uplink and Downlink Communications for Federated Learning
- MURANA: A Generic Framework for Stochastic Variance-Reduced Optimization
- A Robust Gradient Tracking Method for Distributed Optimization over Directed Networks
- Wireless Distributed Edge Learning: How Many Edge Devices Do We Need?
- Optimization for Supervised Machine Learning: Randomized Algorithms for Data and Parameters
- CatFedAvg: Optimising Communication-efficiency and Classification Accuracy in Federated Learning
- Trade-offs of Local SGD at Scale: An Empirical Study
- An Empirical Study on Compressed Decentralized Stochastic Gradient Algorithms with Overparameterized Models
- Toward Efficient Federated Learning in Multi-Channeled Mobile Edge Network with Layerd Gradient Compression
- Toward Communication Efficient Adaptive Gradient Method
- Towards Heterogeneous Clients with Elastic Federated Learning
- Accelerated Stochastic ExtraGradient: Mixing Hessian and Gradient Similarity to Reduce Communication in Distributed and Federated Learning
- Communication-Efficient Federated Linear and Deep Generalized Canonical Correlation Analysis
- Quantizing data for distributed learning
- Det-CGD: Compressed Gradient Descent with Matrix Stepsizes for Non-Convex Optimization
- TOFU: Towards Obfuscated Federated Updates by Encoding Weight Updates into Gradients from Proxy Data
- Communication-Efficient Network-Distributed Optimization with Differential-Coded Compressors
- Towards Sharper First-Order Adversary with Quantized Gradients
- The Gradient Convergence Bound of Federated Multi-Agent Reinforcement Learning with Efficient Communication
- MQGrad: Reinforcement Learning of Gradient Quantization in Parameter Server
- Edge Artificial Intelligence for 6G: Vision, Enabling Technologies, and Applications
- Doing More by Doing Less: How Structured Partial Backpropagation Improves Deep Learning Clusters
- MergeComp: A Compression Scheduler for Scalable Communication-Efficient Distributed Training
- Caramel: Accelerating Decentralized Distributed Deep Learning with Computation Scheduling
- Scalable Projection-Free Optimization
- Progressive Compressed Records: Taking a Byte out of Deep Learning Data
- Masked Training of Neural Networks with Partial Gradients