5 citations · 5 across the 3 of their papers we have counts for
3 papers
cs.CL2025
One Size Does Not Fit All: A Distribution-Aware Sparsification for More Precise Model Merging
Yingfeng Luo, Dingyang Lin, Junxin Wang +8
Model merging has emerged as a compelling data-free paradigm for multi-task learning, enabling the fusion of multiple fine-tuned models into a single, powerful entity. A key techni…
cs.CL2023★ 5 cited
Deliberate then Generate: Enhanced Prompting Framework for Text Generation
Bei Li, Rui Wang, Junliang Guo +7
Large language models (LLMs) have shown remarkable success across a wide range of natural language generation tasks, where proper prompt designs make great impacts. While existing…
cs.CL2023
Improved Knowledge Distillation for Pre-trained Language Models via Knowledge Selection
Chenglong Wang, Yi Lu, Yongyu Mu +3
Knowledge distillation addresses the problem of transferring knowledge from a teacher model to a student model. In this process, we typically have multiple types of knowledge extra…