Showing stat.MLShow all
2 papers · 1 filter
stat.ML2026
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization
Wu Lin, Scott C. Lowe, Felix Dangel +3
Shampoo and its efficient variant, SOAP, employ structured second-moment estimations and have shown strong performance for training neural networks (NNs). In practice, however, Sha…
stat.ML2025
Spectral-factorized Positive-definite Curvature Learning for NN Training
Wu Lin, Felix Dangel, Runa Eschenhagen +3
Many training methods, such as Adam(W) and Shampoo, learn a positive-definite curvature matrix and apply an inverse root before preconditioning. Recently, non-diagonal training met…