Asynchronous Stochastic Gradient Descent with Delay Compensation
arXiv:1609.08326
Abstract
With the fast development of deep learning, it has become common to learn big neural networks using massive training data. Asynchronous Stochastic Gradient Descent (ASGD) is widely adopted to fulfill this task for its efficiency, which is, however, known to suffer from the problem of delayed gradients. That is, when a local worker adds its gradient to the global model, the global model may have been updated by other workers and this gradient becomes "delayed". We propose a novel technology to compensate this delay, so as to make the optimization behavior of ASGD closer to that of sequential SGD. This is achieved by leveraging Taylor expansion of the gradient function and efficient approximation to the Hessian matrix of the loss function. We call the new algorithm Delay Compensated ASGD (DC-ASGD). We evaluated the proposed algorithm on CIFAR-10 and ImageNet datasets, and the experimental results demonstrate that DC-ASGD outperforms both synchronous SGD and asynchronous SGD, and nearly approaches the performance of sequential SGD.
20 pages, 5 figures
References in corpus (3)
Cited by in corpus (35)
- Convergence of Edge Computing and Deep Learning: A Comprehensive Survey
- Asynchronous Federated Optimization
- Accelerating Federated Learning over Reliability-Agnostic Clients in Mobile Edge Computing Systems
- Adaptive Federated Learning in Resource Constrained Edge Computing Systems
- Federated Learning with Buffered Asynchronous Aggregation
- Scale out for large minibatch SGD: Residual network training on ImageNet-1K with improved accuracy and reduced time to train
- Guided parallelized stochastic gradient descent for delay compensation
- Asynchronous Federated Learning with Reduced Number of Rounds and with Differential Privacy from Less Aggregated Gaussian Noise
- Communication-Efficient ADMM-based Federated Learning
- Dynamic Mini-batch SGD for Elastic Distributed Training: Learning in the Limbo of Resources
- Taming Momentum in a Distributed Asynchronous Environment
- Gradient Scheduling with Global Momentum for Non-IID Data Distributed Asynchronous Training
- DBS: Dynamic Batch Size For Distributed Deep Neural Network Training
- Noiseless Privacy-Preserving Decentralized Learning
- NeCPD: An Online Tensor Decomposition with Optimal Stochastic Gradient Descent
- Taming Convergence for Asynchronous Stochastic Gradient Descent with Unbounded Delay in Non-Convex Learning
- Communication-Efficient Federated Learning with Compensated Overlap-FedAvg
- ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training
- Towards Efficient Scheduling of Federated Mobile Devices under Computational and Statistical Heterogeneity
- AdaptCL: Efficient Collaborative Learning with Dynamic and Adaptive Pruning
- Hogwild! over Distributed Local Data Sets with Linearly Increasing Mini-Batch Sizes
- The Convergence of Stochastic Gradient Descent in Asynchronous Shared Memory
- A Survey on Large-scale Machine Learning
- Accelerated Sparsified SGD with Error Feedback
- Adaptive Braking for Mitigating Gradient Delay
- Revolutionizing Wireless Networks with Federated Learning: A Comprehensive Review
- Finite-Time Consensus Learning for Decentralized Optimization with Nonlinear Gossiping
- FedProf: Selective Federated Learning with Representation Profiling
- OD-SGD: One-step Delay Stochastic Gradient Descent for Distributed Training
- Gap Aware Mitigation of Gradient Staleness
- Pushing the boundaries of parallel Deep Learning -- A practical approach
- Learning Under Delayed Feedback: Implicitly Adapting to Gradient Delays
- Accumulated Decoupled Learning: Mitigating Gradient Staleness in Inter-Layer Model Parallelization
- Distributed Learning and its Application for Time-Series Prediction
- Zeno++: Robust Fully Asynchronous SGD