10 citations · 22 across the 21 of their papers we have counts for
24 papers
Warmup-Distill: Bridge the Distribution Mismatch between Teacher and Student before Knowledge Distillation
Zengkui Sun, Yijin Liu, Fandong Meng +3
The widespread deployment of Large Language Models (LLMs) is hindered by the high computational demands, making knowledge distillation (KD) crucial for developing compact smaller o…
Enhancing Cross-Tokenizer Knowledge Distillation with Contextual Dynamical Mapping
Yijie Chen, Yijin Liu, Fandong Meng +3
Knowledge Distillation (KD) has emerged as a prominent technique for model compression. However, conventional KD approaches primarily focus on homogeneous architectures with identi…
Textualized Agent-Style Reasoning for Complex Tasks by Multiple Round LLM Generation
Chen Liang, Zhifan Feng, Zihe Liu +4
Chain-of-thought prompting significantly boosts the reasoning ability of large language models but still faces three issues: hallucination problem, restricted interpretability, and…
BJTU-WeChat's Systems for the WMT22 Chat Translation Task
Yunlong Liang, Fandong Meng, Jinan Xu +2
This paper introduces the joint submission of the Beijing Jiaotong University and WeChat AI to the WMT'22 chat translation task for English-German. Based on the Transformer, we app…
Improved Data Augmentation for Translation Suggestion
Hongxiao Zhang, Siyu Lai, Songming Zhang +4
Translation suggestion (TS) models are used to automatically provide alternative suggestions for incorrect spans in sentences generated by machine translation. This paper introduce…
Cross-Align: Modeling Deep Cross-lingual Interactions for Word Alignment
Siyu Lai, Zhen Yang, Fandong Meng +3
Word alignment which aims to extract lexicon translation equivalents between source and target sentences, serves as a fundamental tool for natural language processing. Recent studi…