2 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CL2023
Co-training and Co-distillation for Quality Improvement and Compression of Language Models
Hayeon Lee, Rui Hou, Jongpil Kim +4
Knowledge Distillation (KD) compresses computationally expensive pre-trained language models (PLMs) by transferring their knowledge to smaller models, allowing their use in resourc…
cs.CL2023★ 1 cited
A Study on Knowledge Distillation from Weak Teacher for Scaling Up Pre-trained Language Models
Hayeon Lee, Rui Hou, Jongpil Kim +3
Distillation from Weak Teacher (DWT) is a method of transferring knowledge from a smaller, weaker teacher model to a larger student model to improve its performance. Previous studi…
cs.LG2023★ 2 cited
Meta-prediction Model for Distillation-Aware NAS on Unseen Datasets
Hayeon Lee, Sohyun An, Minseon Kim +1
Distillation-aware Neural Architecture Search (DaNAS) aims to search for an optimal student architecture that obtains the best performance and/or efficiency when distilling the kno…