4 papers · 1 filter
ODE approximation for the Adam algorithm: General and overparametrized setting
Steffen Dereich, Arnulf Jentzen, Sebastian Kassing
The Adam optimizer is currently presumably the most popular optimization method in deep learning. In this article we develop an ODE based method to study the Adam optimizer in a fa…
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
Steffen Dereich, Robin Graeber, Arnulf Jentzen +1
Gradient descent (GD) based optimization methods are these days the standard tools to train deep neural networks in artificial intelligence systems. In optimization procedures in d…
Sharp higher order convergence rates for the Adam optimizer
Steffen Dereich, Arnulf Jentzen, Adrian Riekert
Gradient descent based optimization methods are the methods of choice to train deep neural networks in machine learning. Beyond the standard gradient descent method, also suitable…
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
Steffen Dereich, Arnulf Jentzen, Adrian Riekert
Deep learning methods - usually consisting of a class of deep neural networks (DNNs) trained by a stochastic gradient descent (SGD) optimization method - are nowadays omnipresent i…