collaborators

6 papers

math.OC2026

On MUON optimization: From non-convergence to an error analysis with Polar Express and the Newton-Schulz polynomial from implementations

Thang Do, Steffen Dereich, Arnulf Jentzen

Stochastic gradient descent (SGD) optimization methods are the standard instruments for the training of deep neural networks (DNNs). In many relevant artificial intelligence (AI) s…

math.OC2026

Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer

Steffen Dereich, Thang Do, Arnulf Jentzen +1

Beside the standard stochastic gradient descent (SGD) method, the Adam optimizer due to Kingma & Ba (2014) is currently probably the best-known optimization method for the training…

cs.LG2026

Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method

Steffen Dereich, Thang Do, Arnulf Jentzen

The adaptive moment estimation (Adam) optimizer proposed by Kingma & Ba (2014) is presumably the most popular stochastic gradient descent (SGD) optimization method for the training…

math.NA2025

Error analysis for the deep Kolmogorov method

Iulian Cîmpean, Thang Do, Lukas Gonon +2

The deep Kolmogorov method is a simple and popular deep learning based method for approximating solutions of partial differential equations (PDEs) of the Kolmogorov type. In this w…

cs.LG2025

Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks

Thang Do, Arnulf Jentzen, Adrian Riekert

Despite the omnipresent use of stochastic gradient descent (SGD) optimization methods in the training of deep neural networks (DNNs), it remains, in basically all practically relev…

cs.LG2025

Mathematical analysis of the gradients in deep learning

Steffen Dereich, Thang Do, Arnulf Jentzen +1

Deep learning algorithms -- typically consisting of a class of deep artificial neural networks (ANNs) trained by a stochastic gradient descent (SGD) optimization method -- are nowa…