1 paper
Guangyu Guo, Dingwen Zhang, Longfei Han +3
Previous knowledge distillation (KD) methods mostly focus on compressing network architectures, which is not thorough enough in deployment as some costs like transmission bandwidth…