activity
20242026
collaborators

11 papers

cs.LG2026

Zeroth-Order Optimization at the Edge of Stability

Minhak Song, Liang Zhang, Bingcong Li +3

Zeroth-order (ZO) methods are widely used when gradients are unavailable or prohibitively expensive, including black-box learning and memory-efficient fine-tuning of large models,…

math.OC2026

Direction-Magnitude Decomposition for Low-Rank Matrix Optimization: Faster Convergence and Saddle-to-saddle Dynamics

Yudong Wei, Liang Zhang, Bingcong Li +1

Low-rank matrix optimization is often carried out via the Burer-Monteiro (BM) formulation, but choosing the factorization rank is delicate and can substantially slow optimizati…

cs.LG2026

On the Benefits of Weight Normalization for Overparameterized Matrix Sensing

Yudong Wei, Liang Zhang, Bingcong Li +1

While normalization techniques are widely used in deep learning, their theoretical understanding remains relatively limited. In this work, we establish the benefits of (generalized…

cs.LG2026

SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model Inference

Hao Ma, Melis Ilayda Bal, Liang Zhang +4

Modern large language models are increasingly deployed under compute and memory constraints, making flexible control of model capacity a central challenge. While sparse and low-ran…

cs.LG2026

Integrated electro-optic attention nonlinearities for transformers

Luis Mickeler, Kai Lion, Alfonso Nardi +5

Transformers have emerged as the dominant neural-network architecture, achieving state-of-the-art performance in language processing and computer vision. At the core of these model…

cs.LG2025

Zeroth-Order Optimization Finds Flat Minima

Liang Zhang, Bingcong Li, Kiran Koshy Thekumparampil +3

Zeroth-order methods are extensively used in machine learning applications where gradients are infeasible or expensive to compute, such as black-box attacks, reinforcement learning…