1 citations · 3 across the 3 of their papers we have counts for
1 paper · 1 filter
Jingyuan Liu, Jianlin Su, Xingcheng Yao +25
Recently, the Muon optimizer based on matrix orthogonalization has demonstrated strong results in training small-scale language models, but the scalability to larger models has not…