GPU Asynchronous Stochastic Gradient Descent to Speed Up Neural Network Training
arXiv:1312.6186
Abstract
The ability to train large-scale neural networks has resulted in state-of-the-art performance in many areas of computer vision. These results have largely come from computational break throughs of two forms: model parallelism, e.g. GPU accelerated training, which has seen quick adoption in computer vision circles, and data parallelism, e.g. A-SGD, whose large scale has been used mostly in industry. We report early experiments with a system that makes use of both model parallelism and data parallelism, we call GPU A-SGD. We show using GPU A-SGD it is possible to speed up training of large convolutional neural networks useful for computer vision. We believe GPU A-SGD will make it possible to train larger networks on larger training sets in a reasonable amount of time.
6 pages, 4 figures
References in corpus (2)
Cited by in corpus (19)
- One weird trick for parallelizing convolutional neural networks
- Deep convolutional neural networks for brain image analysis on magnetic resonance imaging: a review
- Asynchronous Parallel Stochastic Gradient for Nonconvex Optimization
- Recent Advances in Convolutional Neural Networks
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Asynchronous Decentralized Parallel Stochastic Gradient Descent
- Deep learning with Elastic Averaging SGD
- Theano-based Large-Scale Visual Recognition with Multiple GPUs
- Monadic Pavlovian associative learning in a backpropagation-free photonic network
- Communication trade-offs for synchronized distributed SGD with large step size
- Distributed Learning of Deep Neural Networks using Independent Subnet Training
- Optimizing Multi-GPU Parallelization Strategies for Deep Learning Training
- Anarchic Federated Learning
- A Practical Layer-Parallel Training Algorithm for Residual Networks
- Empirical Evaluation of Parallel Training Algorithms on Acoustic Modeling
- Distributed stochastic optimization for deep learning (thesis)
- Deep Learning At Scale and At Ease
- Pushing the boundaries of parallel Deep Learning -- A practical approach
- Distributed stochastic optimization with large delays