4 papers
MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation
Wei Shen, Zhang Yaxiang, Minhui Huang +3
With increasing size of large language models (LLMs), full-parameter fine-tuning imposes substantial memory demands. To alleviate this, we propose a novel memory-efficient training…
On the Convergence Analysis of Muon
Wei Shen, Ruichuan Huang, Minhui Huang +2
The majority of parameters in neural networks are naturally represented as matrices. However, most commonly used optimizers treat these matrix parameters as flattened vectors durin…
A Single-Loop First-Order Algorithm for Linearly Constrained Bilevel Optimization
Wei Shen, Jiawei Zhang, Minhui Huang +1
We study bilevel optimization problems where the lower-level problems are strongly convex and have coupled linear constraints. To overcome the potential non-smoothness of the hyper…
On the Training Convergence of Transformers for In-Context Classification of Gaussian Mixtures
Wei Shen, Ruida Zhou, Jing Yang +1
Although transformers have demonstrated impressive capabilities for in-context learning (ICL) in practice, theoretical understanding of the underlying mechanism that allows transfo…