4 papers
MLorc: Momentum Low-rank Compression for Memory Efficient Large Language Model Adaptation
Wei Shen, Zhang Yaxiang, Minhui Huang +3
With increasing size of large language models (LLMs), full-parameter fine-tuning imposes substantial memory demands. To alleviate this, we propose a novel memory-efficient training…
On the Convergence Analysis of Muon
Wei Shen, Ruichuan Huang, Minhui Huang +2
The majority of parameters in neural networks are naturally represented as matrices. However, most commonly used optimizers treat these matrix parameters as flattened vectors durin…
Feed m Birds with One Scone: Accelerating Multi-task Gradient Balancing via Bi-level Optimization
Xuxing Chen, Yun He, Jiayi Xu +9
In machine learning, the goal of multi-task learning (MTL) is to optimize multiple objectives together. Recent works, for example, Multiple Gradient Descent Algorithm (MGDA) and it…
A Single-Loop First-Order Algorithm for Linearly Constrained Bilevel Optimization
Wei Shen, Jiawei Zhang, Minhui Huang +1
We study bilevel optimization problems where the lower-level problems are strongly convex and have coupled linear constraints. To overcome the potential non-smoothness of the hyper…