Showing math.OCShow all
3 papers · 1 filter
math.OC2026
Free Heavy-Tailed Lunch for Muon: A Theoretical Justification of Empirical Success
Florian Hübler, Thomas Pethick, Suvrit Sra
Non-Euclidean optimisation methods with matrix-valued updates, such as Muon and Scion, have recently shown strong empirical performance for training Transformer models, yet their t…
math.OC2025
Can SGD Handle Heavy-Tailed Noise?
Ilyas Fatkhullin, Florian Hübler, Guanghui Lan
Stochastic Gradient Descent (SGD) is a cornerstone of large-scale optimization, yet its theoretical behavior under heavy-tailed noise -- common in modern machine learning and reinf…
math.OC2025
From Gradient Clipping to Normalization for Heavy Tailed SGD
Florian Hübler, Ilyas Fatkhullin, Niao He
Recent empirical evidence indicates that many machine learning applications involve heavy-tailed gradient noise, which challenges the standard assumptions of bounded variance in st…