6 papers
On MUON optimization: From non-convergence to an error analysis with Polar Express and the Newton-Schulz polynomial from implementations
Thang Do, Steffen Dereich, Arnulf Jentzen
Stochastic gradient descent (SGD) optimization methods are the standard instruments for the training of deep neural networks (DNNs). In many relevant artificial intelligence (AI) s…
Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
Steffen Dereich, Thang Do, Arnulf Jentzen +1
Beside the standard stochastic gradient descent (SGD) method, the Adam optimizer due to Kingma & Ba (2014) is currently probably the best-known optimization method for the training…
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
Steffen Dereich, Thang Do, Arnulf Jentzen
The adaptive moment estimation (Adam) optimizer proposed by Kingma & Ba (2014) is presumably the most popular stochastic gradient descent (SGD) optimization method for the training…
Error analysis for the deep Kolmogorov method
Iulian Cîmpean, Thang Do, Lukas Gonon +2
The deep Kolmogorov method is a simple and popular deep learning based method for approximating solutions of partial differential equations (PDEs) of the Kolmogorov type. In this w…
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
Thang Do, Arnulf Jentzen, Adrian Riekert
Despite the omnipresent use of stochastic gradient descent (SGD) optimization methods in the training of deep neural networks (DNNs), it remains, in basically all practically relev…
Mathematical analysis of the gradients in deep learning
Steffen Dereich, Thang Do, Arnulf Jentzen +1
Deep learning algorithms -- typically consisting of a class of deep artificial neural networks (ANNs) trained by a stochastic gradient descent (SGD) optimization method -- are nowa…