1 paper
Ming Gao, Yanwu Xu, Hao Zhang
Matrix optimizers such as Muon are attractive for large-scale training because they can improve convergence and token efficiency over coordinate-wise optimizers. Muon does this by…