1 paper
Minsoo Kim, Sihwa Lee, Sukjin Hong +2
Knowledge distillation (KD) has been a ubiquitous method for model compression to strengthen the capability of a lightweight model with the transferred knowledge from the teacher.…