Three Factors Influencing Minima in SGD
arXiv:1711.04623
Abstract
We investigate the dynamical and convergent properties of stochastic gradient descent (SGD) applied to Deep Neural Networks (DNNs). Characterizing the relation between learning rate, batch size and the properties of the final minima, such as width or generalization, remains an open question. In order to tackle this problem we investigate the previously proposed approximation of SGD by a stochastic differential equation (SDE). We theoretically argue that three factors - learning rate, batch size and gradient covariance - influence the minima found by SGD. In particular we find that the ratio of learning rate to batch size is a key determinant of SGD dynamics and of the width of the final minima, and that higher values of the ratio lead to wider minima and often better generalization. We confirm these findings experimentally. Further, we include experiments which show that learning rate schedules can be replaced with batch size schedules and that the ratio of learning rate to batch size is an important factor influencing the memorization process.
First two authors contributed equally. Short version accepted into ICLR workshop. Accepted to Artificial Neural Networks and Machine Learning, ICANN 2018
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Understanding deep learning requires rethinking generalization
- One weird trick for parallelizing convolutional neural networks
- Qualitatively characterizing neural network optimization problems
- Towards Understanding Generalization of Deep Learning: Perspective of Loss Landscapes
Cited by in corpus (126)
- A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay
- Quantum Natural Gradient
- Revisiting Small Batch Training for Deep Neural Networks
- Training Tips for the Transformer Model
- Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation
- Don't Use Large Mini-Batches, Use Local SGD
- Dark Experience for General Continual Learning: a Strong, Simple Baseline
- Learning explanations that are hard to vary
- A Modern Take on the Bias-Variance Tradeoff in Neural Networks
- A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
- Towards Theoretically Understanding Why SGD Generalizes Better Than ADAM in Deep Learning
- Energy-entropy competition and the effectiveness of stochastic gradient descent in machine learning
- Deep Learning Theory Review: An Optimal Control and Dynamical Systems Perspective
- Fluctuation-dissipation relations for stochastic gradient descent
- On the Computational Inefficiency of Large Batch Sizes for Stochastic Gradient Descent
- On the Origin of Implicit Regularization in Stochastic Gradient Descent
- Understanding Batch Normalization
- The Heavy-Tail Phenomenon in SGD
- Stochasticity helps to navigate rough landscapes: comparing gradient-descent-based algorithms in the phase retrieval problem
- Explicit Regularisation in Gaussian Noise Injections
- Federated Learning with Buffered Asynchronous Aggregation
- Mean-field inference methods for neural networks
- Asymmetric Valleys: Beyond Sharp and Flat Local Minima
- The Implicit Regularization of Stochastic Gradient Flow for Least Squares
- The Break-Even Point on Optimization Trajectories of Deep Neural Networks
- The Impact of the Mini-batch Size on the Variance of Gradients in Stochastic Gradient Descent
- The Full Spectrum of Deepnet Hessians at Scale: Dynamics with SGD Training and Sample Size
- The Implicit and Explicit Regularization Effects of Dropout
- Mathematical Models of Overparameterized Neural Networks
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima
- Non-Gaussianity of Stochastic Gradient Noise
- Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability
- Recent advances in deep learning theory
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel
- Hausdorff Dimension, Heavy Tails, and Generalization in Neural Networks
- Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient Noise
- On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks
- Time Matters in Regularizing Deep Networks: Weight Decay and Data Augmentation Affect Early Learning Dynamics, Matter Little Near Convergence
- SAT: Improving Adversarial Training via Curriculum-Based Loss Smoothing
- On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)
- Large Scale Structure of Neural Network Loss Landscapes
- Dynamic Mini-batch SGD for Elastic Distributed Training: Learning in the Limbo of Resources
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning Dynamics
- Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model
- Wide-minima Density Hypothesis and the Explore-Exploit Learning Rate Schedule
- On Learning Rates and Schrödinger Operators
- CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers
- F2A2: Flexible Fully-decentralized Approximate Actor-critic for Cooperative Multi-agent Reinforcement Learning
- Towards Understanding Generalization in Gradient-Based Meta-Learning
- RankNEAT: Outperforming Stochastic Gradient Search in Preference Learning Tasks
- A consensus-based global optimization method for high dimensional machine learning problems
- Positively Scale-Invariant Flatness of ReLU Neural Networks
- An Empirical Study of Large-Batch Stochastic Gradient Descent with Structured Covariance Noise
- First Exit Time Analysis of Stochastic Gradient Descent Under Heavy-Tailed Gradient Noise
- How noise affects the Hessian spectrum in overparameterized neural networks
- Sharp Bounds for Federated Averaging (Local SGD) and Continuous Perspective
- Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
- Large-Scale Deep Learning Optimizations: A Comprehensive Survey
- Positive-Negative Momentum: Manipulating Stochastic Gradient Noise to Improve Generalization
- Asymptotic Analysis via Stochastic Differential Equations of Gradient Descent Algorithms in Statistical and Computational Paradigms
- On the interplay between noise and curvature and its effect on optimization and generalization
- Needles in Haystacks: On Classifying Tiny Objects in Large Images
- Is Support Set Diversity Necessary for Meta-Learning?
- Deep Learning is Singular, and That's Good
- Traces of Class/Cross-Class Structure Pervade Deep Learning Spectra
- Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias
- Artificial Neural Variability for Deep Learning: On Overfitting, Noise Memorization, and Catastrophic Forgetting
- Regularizing Neural Networks via Adversarial Model Perturbation
- Optimizing Multi-GPU Parallelization Strategies for Deep Learning Training
- Hessian based analysis of SGD for Deep Nets: Dynamics and Generalization
- On Accelerating Distributed Convex Optimizations
- A Loss Curvature Perspective on Training Instability in Deep Learning
- AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the Fly
- Bidirectional Context-Aware Hierarchical Attention Network for Document Understanding
- On generalization bounds for deep networks based on loss surface implicit regularization
- Retrosynthesis with Attention-Based NMT Model and Chemical Analysis of the "Wrong" Predictions
- Parameter Re-Initialization through Cyclical Batch Size Schedules
- Unifying Regularisation Methods for Continual Learning
- Drawing Multiple Augmentation Samples Per Image During Training Efficiently Decreases Test Error
- The Limiting Dynamics of SGD: Modified Loss, Phase Space Oscillations, and Anomalous Diffusion
- SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize Criticality
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of Symmetry
- Large Learning Rate Tames Homogeneity: Convergence and Balancing Effect
- Improving the convergence of SGD through adaptive batch sizes
- Improved generalization by noise enhancement
- Robustness, Privacy, and Generalization of Adversarial Training
- Critical Learning Periods in Federated Learning
- Stochastic Resetting Mitigates Latent Gradient Bias of SGD from Label Noise
- Fast, Better Training Trick -- Random Gradient
- What Happens after SGD Reaches Zero Loss? --A Mathematical Framework
- Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of Stochasticity
- Noise and Fluctuation of Finite Learning Rate Stochastic Gradient Descent
- Hyperplane Arrangements of Trained ConvNets Are Biased
- Stochasticity of Deterministic Gradient Descent: Large Learning Rate for Multiscale Objective Function
- What training reveals about neural network complexity
- On the Generalization of Models Trained with SGD: Information-Theoretic Bounds and Implications
- Trap of Feature Diversity in the Learning of MLPs
- Subaging in underparametrized Deep Neural Networks
- On the Bias-Variance Tradeoff: Textbooks Need an Update
- Smoothness Analysis of Adversarial Training
- The sharp, the flat and the shallow: Can weakly interacting agents learn to escape bad minima?
- Asymmetric Heavy Tails and Implicit Bias in Gaussian Noise Injections
- Understanding Short-Range Memory Effects in Deep Neural Networks
- Imitating Deep Learning Dynamics via Locally Elastic Stochastic Differential Equations
- Generalisation in fully-connected neural networks for time series forecasting
- Unique Properties of Flat Minima in Deep Networks
- Enhance Diffusion to Improve Robust Generalization
- SQWA: Stochastic Quantized Weight Averaging for Improving the Generalization Capability of Low-Precision Deep Neural Networks
- Pushing the boundaries of parallel Deep Learning -- A practical approach
- How do SGD hyperparameters in natural training affect adversarial robustness?
- Loss Landscape Dependent Self-Adjusting Learning Rates in Decentralized Stochastic Gradient Descent
- A Resizable Mini-batch Gradient Descent based on a Multi-Armed Bandit
- Efficient Classification of Very Large Images with Tiny Objects
- What can linear interpolation of neural network loss landscapes tell us?
- Adaptive norms for deep learning with regularized Newton methods
- On Large Batch Training and Sharp Minima: A Fokker-Planck Perspective
- LRTuner: A Learning Rate Tuner for Deep Neural Networks
- Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II
- Continuous-time Models for Stochastic Optimization Algorithms
- BN-invariant sharpness regularizes the training model to better generalization
- Backtracking gradient descent method for general functions, with applications to Deep Learning
- Inherent Noise in Gradient Based Methods
- Analytic expressions for the output evolution of a deep neural network
- A Sharp Convergence Rate for the Asynchronous Stochastic Gradient Descent
- Relative Flatness and Generalization