12 citations · 34 across the 13 of their papers we have counts for
7 papers
Improving Knowledge Distillation via Regularizing Feature Norm and Direction
Yuzhu Wang, Lechao Cheng, Manni Duan +3
Knowledge distillation (KD) exploits a large well-trained model (i.e., teacher) to train a small student model on the same dataset for the same task. Treating teacher features as k…
ViT-Calibrator: Decision Stream Calibration for Vision Transformer
Lin Chen, Zhijie Jia, Tian Qiu +4
A surge of interest has emerged in utilizing Transformers in diverse vision tasks owing to its formidable performance. However, existing approaches primarily focus on optimizing in…
Life Regression based Patch Slimming for Vision Transformers
Jiawei Chen, Lin Chen, Jiang Yang +4
Vision transformers have achieved remarkable success in computer vision tasks by using multi-head self-attention modules to capture long-range dependencies within images. However,…
Generalization Matters: Loss Minima Flattening via Parameter Hybridization for Efficient Online Knowledge Distillation
Tianli Zhang, Mengqi Xue, Jiangtao Zhang +5
Most existing online knowledge distillation(OKD) techniques typically require sophisticated modules to produce diverse knowledge for improving students' generalization ability. In…
Dual Domain-Adversarial Learning for Audio-Visual Saliency Prediction
Yingzi Fan, Longfei Han, Yue Zhang +3
Both visual and auditory information are valuable to determine the salient regions in videos. Deep convolution neural networks (CNN) showcase strong capacity in coping with the aud…
Long-term Leap Attention, Short-term Periodic Shift for Video Classification
Hao Zhang, Lechao Cheng, Yanbin Hao +1
Video transformer naturally incurs a heavier computation burden than a static vision transformer, as the former processes times longer sequence than the latter under the curren…