4 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.CL2023★ 4 cited
f-Divergence Minimization for Sequence-Level Knowledge Distillation
Yuqiao Wen, Zichao Li, Wenyu Du +1
Knowledge distillation (KD) is the process of transferring knowledge from a large model to a small one. It has gained increasing attention in the natural language processing commun…
cs.CV2023★ 1 cited
Tied-Augment: Controlling Representation Similarity Improves Data Augmentation
Emirhan Kurtulus, Zichao Li, Yann Dauphin +1
Data augmentation methods have played an important role in the recent advance of deep learning models, and have become an indispensable component of state-of-the-art models in semi…