Showing math.OCShow all
2 papers · 1 filter
math.OC2026
On MUON optimization: From non-convergence to an error analysis with Polar Express and the Newton-Schulz polynomial from implementations
Thang Do, Steffen Dereich, Arnulf Jentzen
Stochastic gradient descent (SGD) optimization methods are the standard instruments for the training of deep neural networks (DNNs). In many relevant artificial intelligence (AI) s…
math.OC2026
Adam symmetry theorem: characterization of the convergence of the stochastic Adam optimizer
Steffen Dereich, Thang Do, Arnulf Jentzen +1
Beside the standard stochastic gradient descent (SGD) method, the Adam optimizer due to Kingma & Ba (2014) is currently probably the best-known optimization method for the training…