2 papers
cs.LG2026
MARS-M: When Variance Reduction Meets Matrices
Yifeng Liu, Angela Yuan, Quanquan Gu
Matrix-based preconditioned optimizers, such as Muon, have recently been shown to be more efficient than scalar-based optimizers for training large-scale neural networks, including…
cs.LG2025
Robust Layerwise Scaling Rules by Proper Weight Decay Tuning
Zhiyuan Fan, Yifeng Liu, Qingyue Zhao +2
Empirical scaling laws prescribe how to allocate parameters, data, and compute, while maximal-update parameterization (P) enables learning-rate transfer across widths by equali…