2 papers
cs.CV2025
Perspective-Aware Teaching: Adapting Knowledge for Heterogeneous Distillation
Jhe-Hao Lin, Yi Yao, Chan-Feng Hsu +3
Knowledge distillation (KD) involves transferring knowledge from a pre-trained heavy teacher model to a lighter student model, thereby reducing the inference cost while maintaining…
cs.LG2025
Swapped Logit Distillation via Bi-level Teacher Alignment
Stephen Ekaputra Limantoro, Jhe-Hao Lin, Chih-Yu Wang +4
Knowledge distillation (KD) compresses the network capacity by transferring knowledge from a large (teacher) network to a smaller one (student). It has been mainstream that the tea…