Strong error analysis for stochastic gradient descent optimization algorithms
arXiv:1801.09324 · doi:10.1093/imanum/drz055
Abstract
Stochastic gradient descent (SGD) optimization algorithms are key ingredients in a series of machine learning applications. In this article we perform a rigorous strong error analysis for SGD optimization algorithms. In particular, we prove for every arbitrarily small and every arbitrarily large that the considered SGD optimization algorithm converges in the strong -sense with order to the global minimum of the objective function of the considered stochastic approximation problem under standard convexity-type assumptions on the objective function and relaxed assumptions on the moments of the stochastic errors appearing in the employed SGD optimization algorithm. The key ideas in our convergence proof are, first, to employ techniques from the theory of Lyapunov-type functions for dynamical systems to develop a general convergence machinery for SGD optimization algorithms based on such functions, then, to apply this general machinery to concrete Lyapunov-type functions with polynomial structures, and, thereafter, to perform an induction argument along the powers appearing in the Lyapunov-type functions in order to achieve for every arbitrarily large strong -convergence rates. This article also contains an extensive review of results on SGD optimization algorithms in the scientific literature.
References in corpus (24)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- ADADELTA: An Adaptive Learning Rate Method
- DGM: A deep learning algorithm for solving partial differential equations
- Solving high-dimensional partial differential equations using deep learning
- Proceedings of the 29th International Conference on Machine Learning (ICML-12)
- Convolutional Neural Network Architectures for Matching Natural Language Sentences
- Deep learning-based numerical methods for high-dimensional parabolic partial differential equations and backward stochastic differential equations
- A Convergence Theory for Deep Learning via Over-Parameterization
- A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Machine learning approximation algorithms for high-dimensional fully nonlinear partial differential equations and second-order backward stochastic differential equations
- No More Pesky Learning Rates
- Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured Data
- Sparse Online Learning via Truncated Gradient
- Deep splitting method for parabolic PDEs
- Pricing and hedging American-style options with deep learning
- Deep calibration of rough stochastic volatility models
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- Lower error bounds for the stochastic gradient descent optimization algorithm: Sharp convergence rates for slowly and fast decaying learning rates
- Strong convergence for explicit space-time discrete numerical approximation methods for stochastic Burgers equations
- On stochastic gradient Langevin dynamics with dependent data streams: the fully non-convex case
- When Does Stochastic Gradient Algorithm Work Well?
- AdaBatch: Efficient Gradient Aggregation Rules for Sequential and Parallel Stochastic Gradient Methods
Cited by in corpus (15)
- Machine learning approximation algorithms for high-dimensional fully nonlinear partial differential equations and second-order backward stochastic differential equations
- The Modern Mathematics of Deep Learning
- Full error analysis for the training of deep neural networks
- Analysis of Stochastic Gradient Descent in Continuous Time
- Lower error bounds for the stochastic gradient descent optimization algorithm: Sharp convergence rates for slowly and fast decaying learning rates
- A proof of convergence for gradient descent in the training of artificial neural networks for constant target functions
- On the existence of global minima and convergence analyses for gradient descent methods in the training of deep neural networks
- A proof of convergence for stochastic gradient descent in the training of artificial neural networks with ReLU activation for constant target functions
- Existence, uniqueness, and convergence rates for gradient flows in the training of artificial neural networks with ReLU activation
- Uniform-in-Time Weak Error Analysis for Stochastic Gradient Descent Algorithms via Diffusion Approximation
- High-dimensional approximation spaces of artificial neural networks and applications to partial differential equations
- A Hybrid Two-level MCMC Framework to Accelerate Posterior Mean Estimation with Deep Learning Surrogates for Bayesian Inverse Problems
- A proof of convergence for the gradient descent optimization method with random initializations in the training of neural networks with ReLU activation for piecewise linear target functions
- The Deep Parametric PDE Method: Application to Option Pricing
- Multilevel Monte Carlo learning