4 papers
Free Heavy-Tailed Lunch for Muon: A Theoretical Justification of Empirical Success
Florian Hübler, Thomas Pethick, Suvrit Sra
Non-Euclidean optimisation methods with matrix-valued updates, such as Muon and Scion, have recently shown strong empirical performance for training Transformer models, yet their t…
Muown: Row-Norm Control for Muon Optimization
Kai Lion, Florian Hübler, Bingcong Li +2
Muon has emerged as a strong competitor to AdamW for language model pre-training, yet its behavior at scale is sensitive to weight decay. Recent work has observed that, for Muon wi…
Can SGD Handle Heavy-Tailed Noise?
Ilyas Fatkhullin, Florian Hübler, Guanghui Lan
Stochastic Gradient Descent (SGD) is a cornerstone of large-scale optimization, yet its theoretical behavior under heavy-tailed noise -- common in modern machine learning and reinf…
From Gradient Clipping to Normalization for Heavy Tailed SGD
Florian Hübler, Ilyas Fatkhullin, Niao He
Recent empirical evidence indicates that many machine learning applications involve heavy-tailed gradient noise, which challenges the standard assumptions of bounded variance in st…