15 papers
Zeroth-Order Optimization at the Edge of Stability
Minhak Song, Liang Zhang, Bingcong Li +3
Zeroth-order (ZO) methods are widely used when gradients are unavailable or prohibitively expensive, including black-box learning and memory-efficient fine-tuning of large models,…
Direction-Magnitude Decomposition for Low-Rank Matrix Optimization: Faster Convergence and Saddle-to-saddle Dynamics
Yudong Wei, Liang Zhang, Bingcong Li +1
Low-rank matrix optimization is often carried out via the Burer-Monteiro (BM) formulation, but choosing the factorization rank is delicate and can substantially slow optimizati…
Explainable AI for Next-Generation Wireless Physical Layer: Basics, State-of-the-Art, and Open Challenges
Bingnan Xiao, Shuyan Hu, Xiaojing Chen +5
Next-generation wireless systems are expected to be ``AI-native," with neural networks (NNs) embedded throughout the physical (PHY) layer protocol stack to improve spectral efficie…
On the Benefits of Weight Normalization for Overparameterized Matrix Sensing
Yudong Wei, Liang Zhang, Bingcong Li +1
While normalization techniques are widely used in deep learning, their theoretical understanding remains relatively limited. In this work, we establish the benefits of (generalized…
SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model Inference
Hao Ma, Melis Ilayda Bal, Liang Zhang +4
Modern large language models are increasingly deployed under compute and memory constraints, making flexible control of model capacity a central challenge. While sparse and low-ran…
Muown: Row-Norm Control for Muon Optimization
Kai Lion, Florian Hübler, Bingcong Li +2
Muon has emerged as a strong competitor to AdamW for language model pre-training, yet its behavior at scale is sensitive to weight decay. Recent work has observed that, for Muon wi…