9 citations · 11 across the 5 of their papers we have counts for
1 paper · 1 filter
Ruijie Zhang, Yequan Zhao, Ziyue Liu +4
Muon has recently emerged as a strong optimizer for large language model pre-training, orthogonalizing the momentum matrix via Newton--Schulz polar iterations. A natural intuition…