1 paper
Tianyu Pang, Yujie Fang, Zihang Liu +4
Muon has recently shown promising results in LLM training. In this work, we study how to further improve Muon. We argue that Muon's orthogonalized update rule suppresses the emerge…