3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CV2023
Cumulative Spatial Knowledge Distillation for Vision Transformers
Borui Zhao, Renjie Song, Jiajun Liang
Distilling knowledge from convolutional neural networks (CNNs) is a double-edged sword for vision transformers (ViTs). It boosts the performance since the image-friendly local-indu…
cs.CV2023
DOT: A Distillation-Oriented Trainer
Borui Zhao, Quan Cui, Renjie Song +1
Knowledge distillation transfers knowledge from a large model to a small one via task and distillation losses. In this paper, we observe a trade-off between task and distillation l…
cs.CV2022★ 3 cited
Efficient One Pass Self-distillation with Zipf's Label Smoothing
Jiajun Liang, Linze Li, Zhaodong Bing +4
Self-distillation exploits non-uniform soft supervision from itself during training and improves performance without any runtime cost. However, the overhead during training is ofte…