3 papers
cs.IR2026
Towards Dynamic Dense Retrieval with Routing Strategy
Zhan Su, Fengran Mo, Jinghan Zhang +4
The \textit{de facto} paradigm for applying dense retrieval (DR) to new tasks involves fine-tuning a pre-trained model for a specific task. However, this paradigm has two significa…
cs.LG2025
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
Feiyang Kang, Yifan Sun, Bingbing Wen +4
Domain reweighting is an emerging research area aimed at adjusting the relative weights of different data sources to improve the effectiveness and efficiency of LLM pre-training. W…
cs.LG2025
Tensorized Clustered LoRA Merging for Multi-Task Interference
Zhan Su, Fengran Mo, Guojun Liang +4
Despite the success of the monolithic dense paradigm of large language models (LLMs), the LoRA adapters offer an efficient solution by fine-tuning small task-specific modules and m…