Fluctuation-dissipation relations for stochastic gradient descent
arXiv:1810.00004
Abstract
The notion of the stationary equilibrium ensemble has played a central role in statistical mechanics. In machine learning as well, training serves as generalized equilibration that drives the probability distribution of model parameters toward stationarity. Here, we derive stationary fluctuation-dissipation relations that link measurable quantities and hyperparameters in the stochastic gradient descent algorithm. These relations hold exactly for any stationary state and can in particular be used to adaptively set training schedule. We can further use the relations to efficiently extract information pertaining to a loss-function landscape such as the magnitudes of its Hessian and anharmonicity. Our claims are empirically verified.
15 pages, 6 figures; v2: final version accepted at ICLR 2019, with derivations/assumptions clarified and Adam/AMSGrad experiments added
References in corpus (2)
Cited by in corpus (15)
- How Good is the Bayes Posterior in Deep Neural Networks Really?
- The Full Spectrum of Deepnet Hessians at Scale: Dynamics with SGD Training and Sample Size
- Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
- The Implicit and Explicit Regularization Effects of Dropout
- Shape Matters: Understanding the Implicit Bias of the Noise Covariance
- Active Importance Sampling for Variational Objectives Dominated by Rare Events: Consequences for Optimization and Generalization
- On the interplay between noise and curvature and its effect on optimization and generalization
- Strength of Minibatch Noise in SGD
- Fluctuation-dissipation Type Theorem in Stochastic Linear Learning
- Statistical Adaptive Stochastic Gradient Methods
- Robust Learning Rate Selection for Stochastic Optimization via Splitting Diagnostic
- The Limiting Dynamics of SGD: Modified Loss, Phase Space Oscillations, and Anomalous Diffusion
- SDE approximations of GANs training and its long-run behavior
- Generative Adversarial Network: Some Analytical Perspectives
- Multilayer Lookahead: a Nested Version of Lookahead