Don't Use Large Mini-Batches, Use Local SGD
arXiv:1808.07217
Abstract
Mini-batch stochastic gradient methods (SGD) are state of the art for distributed training of deep neural networks. Drastic increases in the mini-batch sizes have lead to key efficiency and scalability gains in recent years. However, progress faces a major roadblock, as models trained with large batches often do not generalize well, i.e. they do not show good accuracy on new data. As a remedy, we propose a \emph{post-local} SGD and show that it significantly improves the generalization performance compared to large-batch training on standard benchmarks while enjoying the same efficiency (time-to-accuracy) and scalability. We further provide an extensive study of the communication efficiency vs. performance trade-offs associated with a host of \emph{local SGD} variants.
To appear in ICLR 2020
References in corpus (19)
- Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
- Revisiting Distributed Synchronous SGD
- Don't Decay the Learning Rate, Increase the Batch Size
- Large Batch Training of Convolutional Networks
- Revisiting Small Batch Training for Deep Neural Networks
- Three Factors Influencing Minima in SGD
- Qualitatively characterizing neural network optimization problems
- Error Feedback Fixes SignSGD and other Gradient Compression Schemes
- Measuring the Effects of Data Parallelism on Neural Network Training
- An Empirical Model of Large-Batch Training
- Hessian-based Analysis of Large Batch Training and Robustness to Adversaries
- The Anisotropic Noise in Stochastic Gradient Descent: Its Behavior of Escaping from Sharp Minima and Regularization Effects
- Parallel SGD: When does averaging help?
- A Walk with SGD
- On the Computational Inefficiency of Large Batch Sizes for Stochastic Gradient Descent
- Parle: parallelizing stochastic gradient descent
- Communication trade-offs for synchronized distributed SGD with large step size
- On the Ineffectiveness of Variance Reduced Optimization for Deep Learning
- An Empirical Study of Large-Batch Stochastic Gradient Descent with Structured Covariance Noise
Cited by in corpus (100)
- On the Convergence of FedAvg on Non-IID Data
- Machine Learning at the Wireless Edge: Distributed Stochastic Gradient Descent Over-the-Air
- Ensemble Distillation for Robust Model Fusion in Federated Learning
- Federated Optimization in Heterogeneous Networks
- Personalized Federated Learning: A Meta-Learning Approach
- FedPAQ: A Communication-Efficient Federated Learning Method with Periodic Averaging and Quantization
- Federated Meta-Learning with Fast Convergence and Efficient Communication
- Personalized Federated Learning with Moreau Envelopes
- Local SGD Converges Fast and Communicates Little
- On the Convergence of Local Descent Methods in Federated Learning
- A Unified Theory of Decentralized SGD with Changing Topology and Local Updates
- Cooperative SGD: A unified Framework for the Design and Analysis of Communication-Efficient SGD Algorithms
- Measuring the Effects of Data Parallelism on Neural Network Training
- Federated Learning in Mobile Edge Networks: A Comprehensive Survey
- Optimal Client Sampling for Federated Learning
- Variance Reduced Local SGD with Lower Communication Complexity
- Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
- Fast Federated Learning by Balancing Communication Trade-Offs
- The Error-Feedback Framework: Better Rates for SGD with Delayed Gradients and Compressed Communication
- SlowMo: Improving Communication-Efficient Distributed SGD with Slow Momentum
- Central Server Free Federated Learning over Single-sided Trust Social Networks
- Oort: Efficient Federated Learning via Guided Participant Selection
- Minibatch vs Local SGD for Heterogeneous Distributed Learning
- Is Local SGD Better than Minibatch SGD?
- Tighter Theory for Local SGD on Identical and Heterogeneous Data
- Federated Learning with Buffered Asynchronous Aggregation
- Faster On-Device Training Using New Federated Momentum Algorithm
- Personalized Federated Learning using Hypernetworks
- Adaptive Communication Strategies to Achieve the Best Error-Runtime Trade-off in Local-Update SGD
- Federated Learning With Quantized Global Model Updates
- Quasi-Global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous Data
- Decentralized Deep Learning with Arbitrary Communication Compression
- SpreadGNN: Serverless Multi-task Federated Learning for Graph Neural Networks
- Communication-Efficient Local Decentralized SGD Methods
- Federated Learning with Compression: Unified Analysis and Sharp Guarantees
- TornadoAggregate: Accurate and Scalable Federated Learning via the Ring-Based Architecture
- FedGEMS: Federated Learning of Larger Server Models via Selective Knowledge Fusion
- Decentralized Bayesian Learning over Graphs
- Consensus Control for Decentralized Deep Learning
- Local AdaAlter: Communication-Efficient Stochastic Gradient Descent with Adaptive Learning Rates
- STEM: A Stochastic Two-Sided Momentum Algorithm Achieving Near-Optimal Sample and Communication Complexities for Federated Learning
- A Better Alternative to Error Feedback for Communication-Efficient Distributed Learning
- Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks
- BROADCAST: Reducing Both Stochastic and Compression Noise to Robustify Communication-Efficient Federated Learning
- Communication-Efficient Distributed Stochastic AUC Maximization with Deep Neural Networks
- An Accelerated Decentralized Stochastic Proximal Algorithm for Finite Sums
- Achieving Linear Speedup with Partial Worker Participation in Non-IID Federated Learning
- Sparse Communication for Training Deep Networks
- Stochastic Weight Averaging in Parallel: Large-Batch Training that Generalizes Well
- Accurate and Fast Federated Learning via IID and Communication-Aware Grouping
- LASG: Lazily Aggregated Stochastic Gradients for Communication-Efficient Distributed Learning
- Distributed Learning of Deep Neural Networks using Independent Subnet Training
- Bias-Variance Reduced Local SGD for Less Heterogeneous Federated Learning
- Exploiting Unlabeled Data in Smart Cities using Federated Learning
- Parallel Restarted SPIDER -- Communication Efficient Distributed Nonconvex Optimization with Optimal Computation Complexity
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- Local SGD With a Communication Overhead Depending Only on the Number of Workers
- FedDR -- Randomized Douglas-Rachford Splitting Algorithms for Nonconvex Federated Composite Optimization
- Gradient Descent with Compressed Iterates
- Local Stochastic Gradient Descent Ascent: Convergence Analysis and Communication Efficiency
- Sharp Bounds for Federated Averaging (Local SGD) and Continuous Perspective
- Local SGD: Unified Theory and New Efficient Methods
- A Distributed Hierarchical SGD Algorithm with Sparse Global Reduction
- RingFed: Reducing Communication Costs in Federated Learning on Non-IID Data
- Hierarchical Weight Averaging for Deep Neural Networks
- Communication-efficient SGD: From Local SGD to One-Shot Averaging
- Elastic Consistency: A General Consistency Model for Distributed Stochastic Gradient Descent
- MixML: A Unified Analysis of Weakly Consistent Parallel Learning
- On the Convergence of Nested Decentralized Gradient Methods with Multiple Consensus and Gradient Steps
- Accelerating Asynchronous Stochastic Gradient Descent for Neural Machine Translation
- Understanding the Effects of Data Parallelism and Sparsity on Neural Network Training
- Accelerating Gossip SGD with Periodic Global Averaging
- Citadel: Protecting Data Privacy and Model Confidentiality for Collaborative Learning with SGX
- Multi-Level Local SGD for Heterogeneous Hierarchical Networks
- Moshpit SGD: Communication-Efficient Decentralized Training on Heterogeneous Unreliable Devices
- Scalable and Practical Natural Gradient for Large-Scale Deep Learning
- Extrapolation for Large-batch Training in Deep Learning
- Optimal Complexity in Decentralized Training
- Delayed Projection Techniques for Linearly Constrained Problems: Convergence Rates, Acceleration, and Applications
- The Minimax Complexity of Distributed Optimization
- Sparsification as a Remedy for Staleness in Distributed Asynchronous SGD
- CSER: Communication-efficient SGD with Error Reset
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
- Statistical Estimation and Inference via Local SGD in Federated Learning
- Communication-Efficient Distributed Learning via Sparse and Adaptive Stochastic Gradient
- Decentralized Learning with Lazy and Approximate Dual Gradients
- CFedAvg: Achieving Efficient Communication and Fast Convergence in Non-IID Federated Learning
- DaSGD: Squeezing SGD Parallelization Performance in Distributed Training Using Delayed Averaging
- Elastic CoCoA: Scaling In to Improve Convergence
- Shuffle-Exchange Brings Faster: Reduce the Idle Time During Communication for Decentralized Neural Network Training
- Federated Deep AUC Maximization for Heterogeneous Data with a Constant Communication Complexity
- Local AdaGrad-Type Algorithm for Stochastic Convex-Concave Optimization
- OpTorch: Optimized deep learning architectures for resource limited environments
- Towards Heterogeneous Clients with Elastic Federated Learning
- HADFL: Heterogeneity-aware Decentralized Federated Learning Framework
- Communication Efficient Generalized Tensor Factorization for Decentralized Healthcare Networks
- Trade-offs of Local SGD at Scale: An Empirical Study
- Accelerate Distributed Stochastic Descent for Nonconvex Optimization with Momentum
- Addressing Algorithmic Bottlenecks in Elastic Machine Learning with Chicle
- MergeComp: A Compression Scheduler for Scalable Communication-Efficient Distributed Training