Showing math.OCShow all
2 papers · 1 filter
math.OC2026
Convergence Rate Analysis of the AdamW-style Shampoo: Unifying One-Sided and Two-Sided Preconditioning
Huan Li, Yiming Dong, Zhouchen Lin
This paper studies AdamW-style Shampoo, an effective variant of the classical Shampoo that won the external tuning track of the AlgoPerf neural network training competition. Our an…
math.OC2025
On the Convergence Rate of RMSProp and Its Momentum Extension Measured by Norm
Huan Li, Yiming Dong, Zhouchen Lin
Although adaptive gradient methods have been extensively used in deep learning, their convergence rates proved in the literature are all slower than that of SGD, particularly with…