8 papers · 1 filter
AdaMuon: Adaptive Muon Optimizer
Chongjie Si, Debing Zhang, Wei Shen
We propose AdaMuon, a novel optimizer that combines element-wise adaptivity with orthogonal updates for large-scale neural network training. AdaMuon incorporates two tightly couple…
Weight Spectra Induced Efficient Model Adaptation
Chongjie Si, Xuankun Yang, Muqing Liu +5
Large-scale foundation models have demonstrated remarkable versatility across a wide range of downstream tasks. However, fully fine-tuning these models incurs prohibitive computati…
MAP: Revisiting Weight Decomposition for Low-Rank Adaptation
Chongjie Si, Zhiyi Shi, Yadao Wang +3
The rapid development of large language models has revolutionized natural language processing, but their fine-tuning remains computationally expensive, hindering broad deployment.…
Revisiting Sparsity Constraint Under High-Rank Property in Partial Multi-Label Learning
Chongjie Si, Yidan Cui, Fuchao Yang +2
Partial Multi-Label Learning (PML) extends the multi-label learning paradigm to scenarios where each sample is associated with a candidate label set containing both ground-truth la…
Why Can Accurate Models Be Learned from Inaccurate Annotations?
Chongjie Si, Yidan Cui, Fuchao Yang +2
Learning from inaccurate annotations has gained significant attention due to the high cost of precise labeling. However, despite the presence of erroneous labels, models trained on…
NAN: A Training-Free Solution to Coefficient Estimation in Model Merging
Chongjie Si, Kangtao Lv, Jingjing Jiang +6
Model merging offers a training-free alternative to multi-task learning by combining independently fine-tuned models into a unified one without access to raw data. However, existin…