3 papers
stat.ML2026
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions
Katie Everett, Elliot Paquette
Existing theory of momentum assumes that gradients arrive at every parameter at a roughly constant rate, an assumption violated in practice by heavy-tailed data distributions and m…
stat.ML2026
Logarithmic-time Schedules for Scaling Language Models with Momentum
Damien Ferbach, Courtney Paquette, Gauthier Gidel +2
In practice, the hyperparameters and weight-decay in AdamW are typically kept at fixed values. Is there any reason to do otherwise? We show that for large-scale…
stat.ML2025
Dimension-adapted Momentum Outscales SGD
Damien Ferbach, Katie Everett, Gauthier Gidel +2
We investigate scaling laws for stochastic momentum algorithms with small batch on the power law random features model, parameterized by data complexity, target complexity, and mod…