AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks
arXiv:1712.02029
Abstract
Training deep neural networks with Stochastic Gradient Descent, or its variants, requires careful choice of both learning rate and batch size. While smaller batch sizes generally converge in fewer training epochs, larger batch sizes offer more parallelism and hence better computational efficiency. We have developed a new training approach that, rather than statically choosing a single batch size for all epochs, adaptively increases the batch size during the training process. Our method delivers the convergence rate of small batch sizes while achieving performance similar to large batch sizes. We analyse our approach using the standard AlexNet, ResNet, and VGG networks operating on the popular CIFAR-10, CIFAR-100, and ImageNet datasets. Our results demonstrate that learning with adaptive batch sizes can improve performance by factors of up to 6.25 on 4 NVIDIA Tesla P100 GPUs while changing accuracy by less than 1% relative to training with fixed batch sizes.
14 pages
References in corpus (4)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Don't Decay the Learning Rate, Increase the Batch Size
- Feedforward and Recurrent Neural Networks Backward Propagation and Hessian in Matrix Form
Cited by in corpus (34)
- Highly Scalable Deep Learning Training System with Mixed-Precision: Training ImageNet in Four Minutes
- Efficient spectrum prediction and inverse design for plasmonic waveguide systems based on artificial neural networks
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
- Measuring the Effects of Data Parallelism on Neural Network Training
- An Empirical Model of Large-Batch Training
- Review: Deep Learning in Electron Microscopy
- Massively Distributed SGD: ImageNet/ResNet-50 Training in a Flash
- A Progressive Batching L-BFGS Method for Machine Learning
- On the Computational Inefficiency of Large Batch Sizes for Stochastic Gradient Descent
- Semi-Dynamic Load Balancing: Efficient Distributed Learning in Non-Dedicated Environments
- Large batch size training of neural networks with adversarial training and second-order information
- Descending through a Crowded Valley - Benchmarking Deep Learning Optimizers
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning
- Taming Resource Heterogeneity In Distributed ML Training With Dynamic Batching
- On the Utility of Gradient Compression in Distributed Training Systems
- Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks
- Estimation of discrete choice models with hybrid stochastic adaptive batch size algorithms
- Accordion: Adaptive Gradient Communication via Critical Learning Regime Identification
- Stochastic Weight Averaging in Parallel: Large-Batch Training that Generalizes Well
- Dynamic Mini-batch SGD for Elastic Distributed Training: Learning in the Limbo of Resources
- On Periodic Functions as Regularizers for Quantization of Neural Networks
- Large-Batch Training for LSTM and Beyond
- Large-Scale Deep Learning Optimizations: A Comprehensive Survey
- History-Gradient Aided Batch Size Adaptation for Variance Reduced Algorithms
- GeoDA: a geometric framework for black-box adversarial attacks
- Study on the Large Batch Size Training of Neural Networks Based on the Second Order Gradient
- HierTrain: Fast Hierarchical Edge AI Learning with Hybrid Parallelism in Mobile-Edge-Cloud Computing
- Parameter Re-Initialization through Cyclical Batch Size Schedules
- Parallel Complexity of Forward and Backward Propagation
- AccelAT: A Framework for Accelerating the Adversarial Training of Deep Neural Networks through Accuracy Gradient
- Inefficiency of K-FAC for Large Batch Size Training
- Improving the convergence of SGD through adaptive batch sizes
- Concurrent Adversarial Learning for Large-Batch Training
- Fast Jacobian-Vector Product for Deep Networks