On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)
arXiv:2102.12470
Abstract
It is generally recognized that finite learning rate (LR), in contrast to infinitesimal LR, is important for good generalization in real-life deep nets. Most attempted explanations propose approximating finite-LR SGD with Ito Stochastic Differential Equations (SDEs), but formal justification for this approximation (e.g., (Li et al., 2019)) only applies to SGD with tiny LR. Experimental verification of the approximation appears computationally infeasible. The current paper clarifies the picture with the following contributions: (a) An efficient simulation algorithm SVAG that provably converges to the conventionally used Ito SDE approximation. (b) A theoretically motivated testable necessary condition for the SDE approximation and its most famous implication, the linear scaling rule (Goyal et al., 2017), to hold. (c) Experiments using this simulation to demonstrate that the previously proposed SDE approximation can meaningfully capture the training and generalization properties of common deep nets.
36 pages, 20 figures
References in corpus (8)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- One weird trick for parallelizing convolutional neural networks
- Towards Theoretically Understanding Why SGD Generalizes Better Than ADAM in Deep Learning
- A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning Dynamics
- On Learning Rates and Schrödinger Operators
- First Exit Time Analysis of Stochastic Gradient Descent Under Heavy-Tailed Gradient Noise
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning Rate
Cited by in corpus (9)
- On the different regimes of Stochastic Gradient Descent
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning Dynamics
- On Quantum Speedups for Nonconvex Optimization via Quantum Tunneling Walks
- Fractal Structure and Generalization Properties of Stochastic Optimization Algorithms
- Stochastic Training is Not Necessary for Generalization
- The Limiting Dynamics of SGD: Modified Loss, Phase Space Oscillations, and Anomalous Diffusion
- What Happens after SGD Reaches Zero Loss? --A Mathematical Framework
- Imitating Deep Learning Dynamics via Locally Elastic Stochastic Differential Equations
- Noether's Learning Dynamics: Role of Symmetry Breaking in Neural Networks