Yet Another Accelerated SGD: ResNet-50 Training on ImageNet in 74.7 seconds
arXiv:1903.12650
Abstract
There has been a strong demand for algorithms that can execute machine learning as faster as possible and the speed of deep learning has accelerated by 30 times only in the past two years. Distributed deep learning using the large mini-batch is a key technology to address the demand and is a great challenge as it is difficult to achieve high scalability on large clusters without compromising accuracy. In this paper, we introduce optimization methods which we applied to this challenge. We achieved the training time of 74.7 seconds using 2,048 GPUs on ABCI cluster applying these methods. The training throughput is over 1.73 million images/sec and the top-1 validation accuracy is 75.08%.
References in corpus (2)
Cited by in corpus (12)
- Optimization for deep learning: theory and algorithms
- Tesseract: Parallelize the Tensor Parallelism Efficiently
- HPC AI500: The Methodology, Tools, Roofline Performance Models, and Metrics for Benchmarking HPC AI Systems
- An Oracle for Guiding Large-Scale Model/Hybrid Parallel Training of Convolutional Neural Networks
- The Case for Strong Scaling in Deep Learning: Training Large 3D CNNs with Hybrid Parallelism
- Accelerating Neural Network Training with Distributed Asynchronous and Selective Optimization (DASO)
- Stochastic Training is Not Necessary for Generalization
- Crossover-SGD: A gossip-based communication in distributed deep learning for alleviating large mini-batch problem and enhancing scalability
- Concurrent Adversarial Learning for Large-Batch Training
- Scalable and Practical Natural Gradient for Large-Scale Deep Learning
- Scaling Distributed Deep Learning Workloads beyond the Memory Capacity with KARMA
- DEED: A General Quantization Scheme for Communication Efficiency in Bits