20 citations · 23 across the 3 of their papers we have counts for
5 papers
Fixing the Teacher-Student Knowledge Discrepancy in Distillation
Jiangfan Han, Mengya Gao, Yujie Wang +3
Training a small student network with the guidance of a larger teacher network is an effective way to promote the performance of the student. Despite the different types, the guide…
Residual Knowledge Distillation
Mengya Gao, Yujun Shen, Quanquan Li +1
Knowledge distillation (KD) is one of the most potent ways for model compression. The key idea is to transfer the knowledge from a deep teacher model (T) to a shallower student (S)…
P2SGrad: Refined Gradients for Optimizing Deep Face Models
Xiao Zhang, Rui Zhao, Junjie Yan +4
Cosine-based softmax losses significantly improve the performance of deep face recognition networks. However, these losses always include sensitive hyper-parameters which can make…
An Embarrassingly Simple Approach for Knowledge Distillation
Mengya Gao, Yujun Shen, Quanquan Li +5
Knowledge Distillation (KD) aims at improving the performance of a low-capacity student model by inheriting knowledge from a high-capacity teacher model. Previous KD methods typica…
A Novel Hybrid Machine Learning Model for Auto-Classification of Retinal Diseases
C. -H. Huck Yang, Jia-Hong Huang, Fangyu Liu +5
Automatic clinical diagnosis of retinal diseases has emerged as a promising approach to facilitate discovery in areas with limited access to specialists. We propose a novel visual-…