3 papers
cs.LG2026
Online Data Selection Is Implicit Alignment
Aoxiong Zeng, Yuxin Yang, Xiangquan Yang
Supervised fine-tuning (SFT) is often treated as a capability-adaptation step, while alignment is attributed to later preference optimization or reinforcement learning. This separa…
cs.LG2026
Disentangling Task Conflicts in Multi-Task LoRA via Orthogonal Gradient Projection
Ziyu Yang, Guibin Chen, Yuxin Yang +2
Multi-Task Learning (MTL) combined with Low-Rank Adaptation (LoRA) has emerged as a promising direction for parameter-efficient deployment of Large Language Models (LLMs). By shari…
cs.LG2026
Towards Specialized Generalists: A Multi-Task MoE-LoRA Framework for Domain-Specific LLM Adaptation
Yuxin Yang, Aoxiong Zeng, Xiangquan Yang
The rapid evolution of Large Language Models (LLMs) has shifted focus from general-purpose capabilities to domain-specific expertise. However, adapting LLMs to specialized fields s…