11 citations · 15 across the 22 of their papers we have counts for
1 paper · 2 filters
Juno Kim, Eshaan Nichani, Denny Wu +2
Spectral optimizers such as Muon have recently shown strong empirical performance in large-scale language model training, but the source and extent of their advantage remain poorly…