1 paper
Seungu Kang, Songkuk Kim
Knowledge Distillation (KD) is a powerful tool for model compression, yet the precise mechanisms by which student models acquire feature representations remain underexplored. In th…