2 papers
cs.LG2026
Musec: MomentUm SpEctral Clipping for Stable Muon-type Training
Zhuanghua Liu, Menglian Wang, Luo Luo
Muon has emerged as a highly effective optimizer for large language model training, often achieving superior convergence and performance compared with the widely adopted Adam and A…
math.OC2026
Near-Optimal Decentralized Stochastic Nonconvex Optimization with Heavy-Tailed Noise
Menglian Wang, Zhuanghua Liu, Luo Luo
This paper studies decentralized stochastic nonconvex optimization problem over row-stochastic networks. We consider the heavy-tailed gradient noise which is empirically observed i…