5 papers · 1 filter
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
Steffen Dereich, Thang Do, Arnulf Jentzen
The adaptive moment estimation (Adam) optimizer proposed by Kingma & Ba (2014) is presumably the most popular stochastic gradient descent (SGD) optimization method for the training…
SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal Structures
Julian Kranz, Davide Gallon, Steffen Dereich +1
We study gradient flows for loss landscapes of fully connected feedforward neural networks with commonly used continuously differentiable activation functions such as the logistic,…
In almost all shallow analytic neural network optimization landscapes, efficient minimizers have strongly convex neighborhoods
Felix Benning, Steffen Dereich
Whether or not a local minimum of a cost function has a strongly convex neighborhood greatly influences the asymptotic convergence rate of optimizers. In this article, we rigorousl…
Mathematical analysis of the gradients in deep learning
Steffen Dereich, Thang Do, Arnulf Jentzen +1
Deep learning algorithms -- typically consisting of a class of deep artificial neural networks (ANNs) trained by a stochastic gradient descent (SGD) optimization method -- are nowa…
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates
Steffen Dereich, Robin Graeber, Arnulf Jentzen
Deep learning algorithms - typically consisting of a class of deep neural networks trained by a stochastic gradient descent (SGD) optimization method - are nowadays the key ingredi…