3 papers
cs.CL2025
Transferable text data distillation by trajectory matching
Rong Yao, Hailin Hu, Yifei Fu +5
In the realm of large language model (LLM), as the size of large models increases, it also brings higher training costs. There is a urgent need to minimize the data size in LLM tra…
cs.CL2025
Saliency-driven Dynamic Token Pruning for Large Language Models
Yao Tao, Yehui Tang, Yun Wang +3
Despite the recent success of large language models (LLMs), LLMs are particularly challenging in long-sequence inference scenarios due to the quadratic computational complexity of…
cs.CL2023
Data-Free Distillation of Language Model by Text-to-Text Transfer
Zheyuan Bai, Xinduo Liu, Hailin Hu +3
Data-Free Knowledge Distillation (DFKD) plays a vital role in compressing the model when original training data is unavailable. Previous works for DFKD in NLP mainly focus on disti…