4 papers
Optimizer Memory Schedules for Outscaling the Overtraining Axis
Katie Everett, Shikai Qiu
We investigate how optimizers scale across the overtraining axis and show that relative optimizer performance and optimal hyperparameters change substantially with training horizon…
Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions
Katie Everett, Elliot Paquette
Existing theory of momentum assumes that gradients arrive at every parameter at a roughly constant rate, an assumption violated in practice by heavy-tailed data distributions and m…
Logarithmic-time Schedules for Scaling Language Models with Momentum
Damien Ferbach, Courtney Paquette, Gauthier Gidel +2
In practice, the hyperparameters and weight-decay in AdamW are typically kept at fixed values. Is there any reason to do otherwise? We show that for large-scale la…
Dimension-adapted Momentum Outscales SGD
Damien Ferbach, Katie Everett, Gauthier Gidel +2
We investigate scaling laws for stochastic momentum algorithms with small batch on the power law random features model, parameterized by data complexity, target complexity, and mod…