Staleness-aware Async-SGD for Distributed Deep Learning
arXiv:1511.05950
Abstract
Deep neural networks have been shown to achieve state-of-the-art performance in several machine learning tasks. Stochastic Gradient Descent (SGD) is the preferred optimization algorithm for training these networks and asynchronous SGD (ASGD) has been widely adopted for accelerating the training of large-scale deep networks in a distributed computing environment. However, in practice it is quite challenging to tune the training hyperparameters (such as learning rate) when using ASGD so as achieve convergence and linear speedup, since the stability of the optimization algorithm is strongly influenced by the asynchronous nature of parameter updates. In this paper, we propose a variant of the ASGD algorithm in which the learning rate is modulated according to the gradient staleness and provide theoretical guarantees for convergence of this algorithm. Experimental verification is performed on commonly-used image classification benchmarks: CIFAR10 and Imagenet to demonstrate the superior effectiveness of the proposed approach, compared to SSGD (Synchronous SGD) and the conventional ASGD algorithm.
Accepted by IJCAI 2016
References in corpus (1)
Cited by in corpus (41)
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- MD-GAN: Multi-Discriminator Generative Adversarial Networks for Distributed Datasets
- Database Meets Deep Learning: Challenges and Opportunities
- Deep Learning Methods for Solving Linear Inverse Problems: Research Directions and Paradigms
- Towards Efficient and Stable K-Asynchronous Federated Learning with Unbounded Stale Gradients on Non-IID Data
- Efficient and Robust Parallel DNN Training through Model Parallelism on Multi-GPU Platform
- Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
- Hanayo: Harnessing Wave-like Pipeline Parallelism for Enhanced Large Model Training Efficiency
- Training Large Neural Networks with Constant Memory using a New Execution Algorithm
- Accuracy-Efficiency Trade-Offs and Accountability in Distributed ML Systems
- Faster Asynchronous SGD
- Taming Momentum in a Distributed Asynchronous Environment
- A(DP)SGD: Asynchronous Decentralized Parallel Stochastic Gradient Descent with Differential Privacy
- Decentralized Consensus Algorithm with Delayed and Stochastic Gradients
- Accelerating Neural Network Training with Distributed Asynchronous and Selective Optimization (DASO)
- Parareal Neural Networks Emulating a Parallel-in-time Algorithm
- Distributed Asynchronous Dual Free Stochastic Dual Coordinate Ascent
- Adaptive Task Allocation for Asynchronous Federated and Parallelized Mobile Edge Learning
- Elastic Gossip: Distributing Neural Network Training Using Gossip-like Protocols
- MixML: A Unified Analysis of Weakly Consistent Parallel Learning
- Task Allocation for Asynchronous Mobile Edge Learning with Delay and Energy Constraints
- Efficient Training of Convolutional Neural Nets on Large Distributed Systems
- Distributed Deep Learning Strategies For Automatic Speech Recognition
- Device Scheduling and Update Aggregation Policies for Asynchronous Federated Learning
- Consistent Lock-free Parallel Stochastic Gradient Descent for Fast and Stable Convergence
- Map Generation from Large Scale Incomplete and Inaccurate Data Labels
- Making Asynchronous Stochastic Gradient Descent Work for Transformers
- The Convergence of Stochastic Gradient Descent in Asynchronous Shared Memory
- Consensus Based Multi-Layer Perceptrons for Edge Computing
- Adaptive Braking for Mitigating Gradient Delay
- Pushing the boundaries of parallel Deep Learning -- A practical approach
- Improving Efficiency in Large-Scale Decentralized Distributed Training
- Distributed deep learning on edge-devices: feasibility via adaptive compression
- Distributed stochastic optimization for deep learning (thesis)
- SecEL: Privacy-Preserving, Verifiable and Fault-Tolerant Edge Learning for Autonomous Vehicles
- Incentive-based integration of useful work into blockchains
- Towards Understanding Acceleration Tradeoff between Momentum and Asynchrony in Nonconvex Stochastic Optimization
- Approximate Random Dropout
- A Highly Efficient Distributed Deep Learning System For Automatic Speech Recognition
- GaDei: On Scale-up Training As A Service For Deep Learning
- HPSGD: Hierarchical Parallel SGD With Stale Gradients Featuring