Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
Dynamic data sampler for cross-language transfer learning in large language models
Yudong Li, Yuhao Feng, Wen Zhou +4
Large Language Models (LLMs) have gained significant attention in the field of natural language processing (NLP) due to their wide range of applications. However, training LLMs for…
cs.CL2023
Weight-Inherited Distillation for Task-Agnostic BERT Compression
Taiqiang Wu, Cheng Hou, Shanshan Lao +4
Knowledge Distillation (KD) is a predominant approach for BERT compression. Previous KD-based methods focus on designing extra alignment losses for the student model to mimic the b…