3 papers
cs.CL2026
Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training
Shengrui Li, Fei Zhao, Kaiyan Zhao +6
Determining an effective data mixture is a key factor in Large Language Model (LLM) pre-training, where models must balance general competence with proficiency on hard tasks such a…
cs.CL2024
NutePrune: Efficient Progressive Pruning with Numerous Teachers for Large Language Models
Shengrui Li, Junzhe Chen, Xueting Han +1
The considerable size of Large Language Models (LLMs) presents notable deployment challenges, particularly on resource-constrained hardware. Structured pruning, offers an effective…
cs.LG2023
AdapterGNN: Parameter-Efficient Fine-Tuning Improves Generalization in GNNs
Shengrui Li, Xueting Han, Jing Bai
Fine-tuning pre-trained models has recently yielded remarkable performance gains in graph neural networks (GNNs). In addition to pre-training techniques, inspired by the latest wor…