5 papers
When to use what Schatten- norm in deep learning?
Thomas Pethick
Schatten- based optimizers such as Muon have shown promising empirical performance, but there remains seemingly conflicting observations regarding whether they are benefici…
Free Heavy-Tailed Lunch for Muon: A Theoretical Justification of Empirical Success
Florian Hübler, Thomas Pethick, Suvrit Sra
Non-Euclidean optimisation methods with matrix-valued updates, such as Muon and Scion, have recently shown strong empirical performance for training Transformer models, yet their t…
Optimistic Dual Averaging Unifies Modern Optimizers
Thomas Pethick, Wanyun Xie, Roman Machacek +1
We introduce SODA, a generalization of Optimistic Dual Averaging, which provides a common perspective on state-of-the-art optimizers like Muon, Lion, AdEMAMix and NAdam, showing th…
Generalized Gradient Norm Clipping & Non-Euclidean -Smoothness
Thomas Pethick, Wanyun Xie, Mete Erdogan +3
This work introduces a hybrid non-Euclidean optimization method which generalizes gradient norm clipping by combining steepest descent and conditional gradient approaches. The meth…
Training Neural Networks at Any Scale
Thomas Pethick, Kimon Antonakopoulos, Antonio Silveti-Falls +2
This article reviews modern optimization methods for training neural networks with an emphasis on efficiency and scale. We present state-of-the-art optimization algorithms under a…