Don't Decay the Learning Rate, Increase the Batch Size
arXiv:1711.00489
Abstract
It is common practice to decay the learning rate. Here we show one can usually obtain the same learning curve on both training and test sets by instead increasing the batch size during training. This procedure is successful for stochastic gradient descent (SGD), SGD with momentum, Nesterov momentum, and Adam. It reaches equivalent test accuracies after the same number of training epochs, but with fewer parameter updates, leading to greater parallelism and shorter training times. We can further reduce the number of parameter updates by increasing the learning rate and scaling the batch size . Finally, one can increase the momentum coefficient and scale , although this tends to slightly reduce the test accuracy. Crucially, our techniques allow us to repurpose existing training schedules for large batch training with no hyper-parameter tuning. We train ResNet-50 on ImageNet to validation accuracy in under 30 minutes.
11 pages, 8 figures. Published as a conference paper at ICLR 2018
References in corpus (6)
- Understanding deep learning requires rethinking generalization
- One weird trick for parallelizing convolutional neural networks
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Large Batch Training of Convolutional Networks
- Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes
- ImageNet Training in Minutes
Cited by in corpus (169)
- A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay
- Revisiting Small Batch Training for Deep Neural Networks
- Highly Scalable Deep Learning Training System with Mixed-Precision: Training ImageNet in Four Minutes
- Training Tips for the Transformer Model
- GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
- Don't Use Large Mini-Batches, Use Local SGD
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Characterizing Implicit Bias in Terms of Optimization Geometry
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
- Bag of Tricks for Image Classification with Convolutional Neural Networks
- An Empirical Model of Large-Batch Training
- Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural Networks
- Measurement of Anomalous Diffusion Using Recurrent Neural Networks
- AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks
- Linear Mode Connectivity and the Lottery Ticket Hypothesis
- Applying Deep Learning To Airbnb Search
- Understanding the Role of Training Regimes in Continual Learning
- Revisiting Training Strategies and Generalization Performance in Deep Metric Learning
- Deep learning enabled design of complex transmission matrices for universal optical components
- Massively Distributed SGD: ImageNet/ResNet-50 Training in a Flash
- Super-resolution emulator of cosmological simulations using deep physical models
- A Modern Take on the Bias-Variance Tradeoff in Neural Networks
- Communication optimization strategies for distributed deep neural network training: A survey
- A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
- On Characterizing the Capacity of Neural Networks using Algebraic Topology
- A Progressive Batching L-BFGS Method for Machine Learning
- The large learning rate phase of deep learning: the catapult mechanism
- Optimizing Network Performance for Distributed DNN Training on GPU Clusters: ImageNet/AlexNet Training in 1.5 Minutes
- Relevance of Rotationally Equivariant Convolutions for Predicting Molecular Properties
- Discovering Low-Precision Networks Close to Full-Precision Networks for Efficient Embedded Inference
- On the Computational Inefficiency of Large Batch Sizes for Stochastic Gradient Descent
- Understanding Batch Normalization
- Federated Learning with Buffered Asynchronous Aggregation
- Scale out for large minibatch SGD: Residual network training on ImageNet-1K with improved accuracy and reduced time to train
- Asymmetric Valleys: Beyond Sharp and Flat Local Minima
- The Implicit Regularization of Stochastic Gradient Flow for Least Squares
- Implicit Bias of Gradient Descent on Linear Convolutional Networks
- Large batch size training of neural networks with adversarial training and second-order information
- Adaptive Communication Strategies to Achieve the Best Error-Runtime Trade-off in Local-Update SGD
- Small-GAN: Speeding Up GAN Training Using Core-sets
- CASI: A Convolutional Neural Network Approach for Shell Identification
- Baryon acoustic oscillations reconstruction using convolutional neural networks
- Characterization of anomalous diffusion through convolutional transformers
- Dim but not entirely dark: Extracting the Galactic Center Excess' source-count distribution with neural nets
- Towards constraining warm dark matter with stellar streams through neural simulation-based inference
- The DNNLikelihood: enhancing likelihood distribution with Deep Learning
- SmoothOut: Smoothing Out Sharp Minima to Improve Generalization in Deep Learning
- A Closer Look at Deep Policy Gradients
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima
- EdgeNet: Semantic Scene Completion from a Single RGB-D Image
- Network slicing for vehicular communications: a multi-agent deep reinforcement learning approach
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning
- Black-Box Optimization with Local Generative Surrogates
- On the adequacy of untuned warmup for adaptive optimization
- Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient Noise
- TF-Replicator: Distributed Machine Learning for Researchers
- Maximizing Parallelism in Distributed Training for Huge Neural Networks
- On Scale-out Deep Learning Training for Cloud and HPC
- On the Utility of Gradient Compression in Distributed Training Systems
- Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks
- improving partition-block-based acoustic echo canceler in under-modeling scenarios
- Speech-Image Semantic Alignment Does Not Depend on Any Prior Classification Tasks
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse Training
- On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)
- Stochastic Training of Residual Networks: a Differential Equation Viewpoint
- Shape Matters: Understanding the Implicit Bias of the Noise Covariance
- Adaptive Distributed Stochastic Gradient Descent for Minimizing Delay in the Presence of Stragglers
- Dynamic Hard Pruning of Neural Networks at the Edge of the Internet
- Neural Sign Language Translation based on Human Keypoint Estimation
- Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
- Dynamic Mini-batch SGD for Elastic Distributed Training: Learning in the Limbo of Resources
- Optimizer Benchmarking Needs to Account for Hyperparameter Tuning
- Deep learning and high harmonic generation
- Distributed Learning of Deep Neural Networks using Independent Subnet Training
- Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model
- Neural Likelihoods via Cumulative Distribution Functions
- Gradient Amplification: An efficient way to train deep neural networks
- HPC AI500: The Methodology, Tools, Roofline Performance Models, and Metrics for Benchmarking HPC AI Systems
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- Phase transitions in the mini-batch size for sparse and dense two-layer neural networks
- An Empirical Study of Large-Batch Stochastic Gradient Descent with Structured Covariance Noise
- On the use of neural networks for the structural characterization of polymeric porous materials
- Towards recognizing the light facet of the Higgs Boson
- Minimum weight norm models do not always generalize well for over-parameterized problems
- History-Gradient Aided Batch Size Adaptation for Variance Reduced Algorithms
- An Empirical Study on Hyperparameters and their Interdependence for RL Generalization
- Modularity in Deep Learning: A Survey
- Positive-Negative Momentum: Manipulating Stochastic Gradient Noise to Improve Generalization
- Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
- Is SGD a Bayesian sampler? Well, almost
- A reinforcement learning path planning approach for range-only underwater target localization with autonomous vehicles
- Neural-encoding Human Experts' Domain Knowledge to Warm Start Reinforcement Learning
- ExT5: Towards Extreme Multi-Task Scaling for Transfer Learning
- Large-scale Pretraining for Neural Machine Translation with Tens of Billions of Sentence Pairs
- sigmoidF1: A Smooth F1 Score Surrogate Loss for Multilabel Classification
- On stochastic mirror descent with interacting particles: convergence properties and variance reduction
- Robust Learning Rate Selection for Stochastic Optimization via Splitting Diagnostic
- ILASR: Privacy-Preserving Incremental Learning for Automatic Speech Recognition at Production Scale
- Label Noise SGD Provably Prefers Flat Global Minimizers
- BrainSlug: Transparent Acceleration of Deep Learning Through Depth-First Parallelism
- Non-Differentiable Supervised Learning with Evolution Strategies and Hybrid Methods
- Zero-shot Neural Passage Retrieval via Domain-targeted Synthetic Question Generation
- Optimizing Multi-GPU Parallelization Strategies for Deep Learning Training
- Robust Optimization for Multilingual Translation with Imbalanced Data
- Parameter Re-Initialization through Cyclical Batch Size Schedules
- A Multigrid Method for Efficiently Training Video Models
- Variance reduction for Riemannian non-convex optimization with batch size adaptation
- REX: Revisiting Budgeted Training with an Improved Schedule
- Fine-Grained Stochastic Architecture Search
- The Limiting Dynamics of SGD: Modified Loss, Phase Space Oscillations, and Anomalous Diffusion
- Normalization in Training U-Net for 2D Biomedical Semantic Segmentation
- Dynamic Sparse Graph for Efficient Deep Learning
- Stochastic Optimization with Laggard Data Pipelines
- Gradient Energy Matching for Distributed Asynchronous Gradient Descent
- Inefficiency of K-FAC for Large Batch Size Training
- Hydra: A Peer to Peer Distributed Training & Data Collection Framework
- Nonlinear Conjugate Gradients For Scaling Synchronous Distributed DNN Training
- Understanding the Effects of Data Parallelism and Sparsity on Neural Network Training
- WHO 2016 subtyping and automated segmentation of glioma using multi-task deep learning
- Improving the convergence of SGD through adaptive batch sizes
- Fast, Better Training Trick -- Random Gradient
- Character-Level Feature Extraction with Densely Connected Networks
- Why flatness does and does not correlate with generalization for deep neural networks
- Bayesian Cycle-Consistent Generative Adversarial Networks via Marginalizing Latent Sampling
- Adaptive Periodic Averaging: A Practical Approach to Reducing Communication in Distributed Learning
- Neural Machine Translation: A Review and Survey
- Reverse engineering learned optimizers reveals known and novel mechanisms
- Policy Information Capacity: Information-Theoretic Measure for Task Complexity in Deep Reinforcement Learning
- Acceleration via Fractal Learning Rate Schedules
- Making Asynchronous Stochastic Gradient Descent Work for Transformers
- Concurrent Adversarial Learning for Large-Batch Training
- Joint Sampling and Optimisation for Inverse Rendering
- Sync-Switch: Hybrid Parameter Synchronization for Distributed Deep Learning
- Understanding Learning Dynamics for Neural Machine Translation
- Stochastic natural gradient descent draws posterior samples in function space
- AccUDNN: A GPU Memory Efficient Accelerator for Training Ultra-deep Neural Networks
- Machine Translation between Vietnamese and English: an Empirical Study
- Scaling Distributed Training of Flood-Filling Networks on HPC Infrastructure for Brain Mapping
- Trust-Region Algorithms for Training Responses: Machine Learning Methods Using Indefinite Hessian Approximations
- Pushing the boundaries of parallel Deep Learning -- A practical approach
- Adapting Convolutional Neural Networks for Geographical Domain Shift
- Predicting the outputs of finite deep neural networks trained with noisy gradients
- OD-SGD: One-step Delay Stochastic Gradient Descent for Distributed Training
- How do SGD hyperparameters in natural training affect adversarial robustness?
- Gradient-only line searches to automatically determine learning rates for a variety of stochastic training algorithms
- Genetic-algorithm-optimized neural networks for gravitational wave classification
- How Data Augmentation affects Optimization for Linear Regression
- Enhance Diffusion to Improve Robust Generalization
- Primitive Agentic First-Order Optimization
- BFTrainer: Low-Cost Training of Neural Networks on Unfillable Supercomputer Nodes
- Digital video microscopy enhanced by deep learning
- Online Evolutionary Batch Size Orchestration for Scheduling Deep Learning Workloads in GPU Clusters
- On Batch Orthogonalization Layers
- Demystifying the Effects of Non-Independence in Federated Learning
- NeuSE: A Neural Snapshot Ensemble Method for Collaborative Filtering
- Implicit Regularization of Bregman Proximal Point Algorithm and Mirror Descent on Separable Data
- A Resizable Mini-batch Gradient Descent based on a Multi-Armed Bandit
- Batch size-invariance for policy optimization
- Neural Architecture Search in Embedding Space
- Hybrid BYOL-ViT: Efficient approach to deal with small datasets
- On tuning deep learning models: a data mining perspective
- Predicting protein secondary structure with Neural Machine Translation
- Long Short Term Memory Networks for Bandwidth Forecasting in Mobile Broadband Networks under Mobility
- Using a one dimensional parabolic model of the full-batch loss to estimate learning rates during training
- Data optimization for large batch distributed training of deep neural networks
- Addressing Algorithmic Bottlenecks in Elastic Machine Learning with Chicle
- Adaptive Weight Decay for Deep Neural Networks