1 citations · 1 across the 7 of their papers we have counts for
8 papers
Musec: MomentUm SpEctral Clipping for Stable Muon-type Training
Zhuanghua Liu, Menglian Wang, Luo Luo
Muon has emerged as a highly effective optimizer for large language model training, often achieving superior convergence and performance compared with the widely adopted Adam and A…
Zeroth-Order Nonconvex Nonsmooth Optimization with Heavy-Tailed Noise
Zhuanghua Liu, Luo Luo
This paper considers the nonconvex nonsmooth problem in which the objective function is Lipschitz continuous. We focus on the stochastic setting where the algorithm can access stoc…
Near-Optimal Decentralized Stochastic Nonconvex Optimization with Heavy-Tailed Noise
Menglian Wang, Zhuanghua Liu, Luo Luo
This paper studies decentralized stochastic nonconvex optimization problem over row-stochastic networks. We consider the heavy-tailed gradient noise which is empirically observed i…
Stochastic Bilevel Optimization with Heavy-Tailed Noise
Zhuanghua Liu, Luo Luo
This paper considers the smooth bilevel optimization in which the lower-level problem is strongly convex and the upper-level problem is possibly nonconvex. We focus on the stochast…
Incremental Gauss--Newton Methods with Superlinear Convergence Rates
Zhiling Zhou, Zhuanghua Liu, Chengchang Liu +1
This paper addresses the challenge of solving large-scale nonlinear equations with Hölder continuous Jacobians. We introduce a novel Incremental Gauss--Newton (IGN) method within e…
Differentiable Cluster Graph Neural Network
Yanfei Dong, Mohammed Haroon Dupty, Lambert Deng +3
Graph Neural Networks often struggle with long-range information propagation and in the presence of heterophilous neighborhoods. We address both challenges with a unified framework…