1 paper
Jae-woong Lee, Minjin Choi, Jongwuk Lee +1
Knowledge distillation (KD) is a well-known method to reduce inference latency by compressing a cumbersome teacher model to a small student model. Despite the success of KD in the…