4 papers
Online Data Selection Is Implicit Alignment
Aoxiong Zeng, Yuxin Yang, Xiangquan Yang
Supervised fine-tuning (SFT) is often treated as a capability-adaptation step, while alignment is attributed to later preference optimization or reinforcement learning. This separa…
The Long-Term Effects of Data Selection in LLM Fine-Tuning
Yuxin Yang, Aoxiong Zeng, Xiangquan Yang
Data selection is increasingly used to reduce the cost of large language model (LLM) fine-tuning, with recent methods prioritizing samples by current utility, diversity, quality, o…
Disentangling Task Conflicts in Multi-Task LoRA via Orthogonal Gradient Projection
Ziyu Yang, Guibin Chen, Yuxin Yang +2
Multi-Task Learning (MTL) combined with Low-Rank Adaptation (LoRA) has emerged as a promising direction for parameter-efficient deployment of Large Language Models (LLMs). By shari…
Towards Specialized Generalists: A Multi-Task MoE-LoRA Framework for Domain-Specific LLM Adaptation
Yuxin Yang, Aoxiong Zeng, Xiangquan Yang
The rapid evolution of Large Language Models (LLMs) has shifted focus from general-purpose capabilities to domain-specific expertise. However, adapting LLMs to specialized fields s…