paper

Convergence Rate Analysis of the AdamW-style Shampoo: Unifying One-Sided and Two-Sided Preconditioning

arXiv:2601.07326

Abstract

This paper studies AdamW-style Shampoo, an effective variant of the classical Shampoo that won the external tuning track of the AlgoPerf neural network training competition. Our analysis unifies one-sided and two-sided preconditioning. When the exponents of the two preconditioners sum to , we establish the convergence rate , where represents the number of iterations, denotes the dimensions of the matrix-valued parameters, and matches the constant appearing in the optimal convergence rate of SGD. Theoretically, the nuclear norm and Frobenius norm satisfy , which suggests that our convergence rate is analogous to the optimal convergence rate of SGD in the ideal case where and and are of comparable magnitude. Then, we extend our analysis to settings where the preconditioning exponents do not sum to 1/2, and establish convergence with an explicit but more involved rate.

V3:ICML Camera-Ready. V4 v.s. V3: extend to the more general setting where the exponents of the two preconditioners do not sum to 1/2