5 papers
Recurrent Knowledge Identification and Fusion for Language Model Continual Learning
Yujie Feng, Xujia Wang, Zexin Lu +7
Continual learning (CL) is crucial for deploying large language models (LLMs) in dynamic real-world environments without costly retraining. While recent model ensemble and model me…
Understanding Layer Significance in LLM Alignment
Guangyuan Shi, Zexin Lu, Xiaoyu Dong +4
Aligning large language models (LLMs) through supervised fine-tuning is essential for tailoring them to specific applications. Recent studies suggest that alignment primarily adjus…
Effectiveness of Pre-training for Few-shot Intent Classification
Haode Zhang, Yuwei Zhang, Li-Ming Zhan +4
This paper investigates the effectiveness of pre-training for few-shot intent classification. While existing paradigms commonly further pre-train language models such as BERT on a…
TaSL: Continual Dialog State Tracking via Task Skill Localization and Consolidation
Yujie Feng, Xu Chu, Yongxin Xu +3
A practical dialogue system requires the capacity for ongoing skill acquisition and adaptability to new tasks while preserving prior knowledge. However, current methods for Continu…
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs
Xiangyu Zhao, Bo Liu, Qijiong Liu +2
We present EasyGen, an efficient model designed to enhance multimodal understanding and generation by harnessing the capabilities of diffusion models and large language models (LLM…