A Tight Convergence Analysis for Stochastic Gradient Descent with Delayed Updates
arXiv:1806.10188
Abstract
We provide tight finite-time convergence bounds for gradient descent and stochastic gradient descent on quadratic functions, when the gradients are delayed and reflect iterates from rounds ago. First, we show that without stochastic noise, delays strongly affect the attainable optimization error: In fact, the error can be as bad as non-delayed gradient descent ran on only of the gradients. In sharp contrast, we quantify how stochastic noise makes the effect of delays negligible, improving on previous work which only showed this phenomenon asymptotically or for much smaller delays. Also, in the context of distributed optimization, the results indicate that the performance of gradient descent with delays is competitive with synchronous approaches such as mini-batching. Our results are based on a novel technique for analyzing convergence of optimization algorithms using generating functions.
Cited by in corpus (13)
- The Error-Feedback Framework: Better Rates for SGD with Delayed Gradients and Compressed Communication
- Fast Federated Learning in the Presence of Arbitrary Device Unavailability
- Stragglers Are Not Disaster: A Hybrid Federated Learning Algorithm with Delayed Gradients
- IDEAL: Inexact DEcentralized Accelerated Augmented Lagrangian Method
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?
- Moshpit SGD: Communication-Efficient Decentralized Training on Heterogeneous Unreliable Devices
- The Minimax Complexity of Distributed Optimization
- Asynchronous Stochastic Optimization Robust to Arbitrary Delays
- Asynchronous Distributed Optimization with Stochastic Delays
- Critical Parameters for Scalable Distributed Learning with Large Batches and Asynchronous Updates
- Asynchronous Iterations in Optimization: New Sequence Results and Sharper Algorithmic Guarantees
- Masked Training of Neural Networks with Partial Gradients
- Learning Under Delayed Feedback: Implicitly Adapting to Gradient Delays