1 citations · 1 across the 4 of their papers we have counts for
4 papers
Warmup-Distill: Bridge the Distribution Mismatch between Teacher and Student before Knowledge Distillation
Zengkui Sun, Yijin Liu, Fandong Meng +3
The widespread deployment of Large Language Models (LLMs) is hindered by the high computational demands, making knowledge distillation (KD) crucial for developing compact smaller o…
Enhancing Cross-Tokenizer Knowledge Distillation with Contextual Dynamical Mapping
Yijie Chen, Yijin Liu, Fandong Meng +3
Knowledge Distillation (KD) has emerged as a prominent technique for model compression. However, conventional KD approaches primarily focus on homogeneous architectures with identi…
Personalized Language Model Learning on Text Data Without User Identifiers
Yucheng Ding, Yangwenjian Tan, Xiangyu Liu +6
In many practical natural language applications, user data are highly sensitive, requiring anonymous uploads of text data from mobile devices to the cloud without user identifiers.…
MaskMamba: A Hybrid Mamba-Transformer Model for Masked Image Generation
Wenchao Chen, Liqiang Niu, Ziyao Lu +2
Image generation models have encountered challenges related to scalability and quadratic complexity, primarily due to the reliance on Transformer-based backbones. In this study, we…