collaborators

12 papers

math.OC2026

On MUON optimization: From non-convergence to an error analysis with Polar Express and the Newton-Schulz polynomial from implementations

Thang Do, Steffen Dereich, Arnulf Jentzen

Stochastic gradient descent (SGD) optimization methods are the standard instruments for the training of deep neural networks (DNNs). In many relevant artificial intelligence (AI) s…

math.OC2026

Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses

Steffen Dereich, Arnulf Jentzen, Adrian Riekert

The standard stochastic gradient descent (SGD) optimization method, as well as adaptive methods such as the Adam optimizer fail to converge if the learning rates do not converge to…

math.OC2026

Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer

Steffen Dereich, Thang Do, Arnulf Jentzen +1

Beside the standard stochastic gradient descent (SGD) method, the Adam optimizer due to Kingma & Ba (2014) is currently probably the best-known optimization method for the training…

math.PR2026

Central limit theorem for the averaged Adam optimizer

Steffen Dereich, Arnulf Jentzen

In this article, we analyse convergence of the averaged Adam optimizer to an attracting zero of the Adam vector field. We provide a central limit theorem that, in particular, quant…

cs.LG2026

Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method

Steffen Dereich, Thang Do, Arnulf Jentzen

The adaptive moment estimation (Adam) optimizer proposed by Kingma & Ba (2014) is presumably the most popular stochastic gradient descent (SGD) optimization method for the training…

cs.LG2026

SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal Structures

Julian Kranz, Davide Gallon, Steffen Dereich +1

We study gradient flows for loss landscapes of fully connected feedforward neural networks with commonly used continuously differentiable activation functions such as the logistic,…