10 citations · 28 across the 11 of their papers we have counts for
10 papers · 1 filter
Asymmetric Decision-Making in Online Knowledge Distillation:Unifying Consensus and Divergence
Zhaowei Chen, Borui Zhao, Yuchen Ge +3
Online Knowledge Distillation (OKD) methods streamline the distillation training process into a single stage, eliminating the need for knowledge transfer from a pretrained teacher…
Cumulative Spatial Knowledge Distillation for Vision Transformers
Borui Zhao, Renjie Song, Jiajun Liang
Distilling knowledge from convolutional neural networks (CNNs) is a double-edged sword for vision transformers (ViTs). It boosts the performance since the image-friendly local-indu…
DOT: A Distillation-Oriented Trainer
Borui Zhao, Quan Cui, Renjie Song +1
Knowledge distillation transfers knowledge from a large model to a small one via task and distillation losses. In this paper, we observe a trade-off between task and distillation l…
Is Synthetic Data From Diffusion Models Ready for Knowledge Distillation?
Zheng Li, Yuxuan Li, Penghai Zhao +3
Diffusion models have recently achieved astonishing performance in generating high-fidelity photo-realistic images. Given their huge success, it is still unclear whether synthetic…
Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data
Yuhao Chen, Xin Tan, Borui Zhao +4
Semi-supervised learning (SSL) has attracted enormous attention due to its vast potential of mitigating the dependence on large labeled datasets. The latest methods (e.g., FixMatch…
Curriculum Temperature for Knowledge Distillation
Zheng Li, Xiang Li, Lingfeng Yang +5
Most existing distillation methods ignore the flexible role of the temperature in the loss function and fix it as a hyper-parameter that can be decided by an inefficient grid searc…