Convergence of stochastic gradient descent schemes for Lojasiewicz-landscapes
arXiv:2102.09385 · doi:10.4208/jml.240109
Abstract
In this article, we consider convergence of stochastic gradient descent schemes (SGD), including momentum stochastic gradient descent (MSGD), under weak assumptions on the underlying landscape. More explicitly, we show that on the event that the SGD stays bounded we have convergence of the SGD if there is only a countable number of critical points or if the objective function satisfies Lojasiewicz-inequalities around all critical levels as all analytic functions do. In particular, we show that for neural networks with analytic activation function such as softplus, sigmoid and the hyperbolic tangent, SGD converges on the event of staying bounded, if the random variables modelling the signal and response in the training are compactly supported.
References in corpus (11)
- On the Almost Sure Convergence of Stochastic Gradient Descent in Non-Convex Problems
- On the existence of global minima and convergence analyses for gradient descent methods in the training of deep neural networks
- Cooling down stochastic differential equations: almost sure convergence
- Stochastic gradient descent with noise of machine learning type. Part II: Continuous time analysis
- Sharp Analysis of Stochastic Optimization under Global Kurdyka-Łojasiewicz Inequality
- On the existence of optimal shallow feedforward networks with ReLU activation
- A Unified Convergence Theorem for Stochastic Optimization Methods
- Fast convergence to non-isolated minima: four equivalent conditions for functions
- A stochastic use of the Kurdyka-Lojasiewicz property: Investigation of optimization algorithms behaviours in a non-convex differentiable framework
- Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting
- A Normal Map-Based Proximal Stochastic Gradient Method: Convergence and Identification Properties