1 paper
Runa Eschenhagen, Anna Cai, Tsung-Hsien Lee +1
Optimizers leveraging the matrix structure in neural networks, such as Shampoo and Muon, are more data-efficient than element-wise algorithms like Adam and Signum. While in specifi…