3 papers
cs.CV2026
Attention Transfer Is Not Universally Effective for Vision Transformers
Huaiyuan Qin, Muli Yang, Gabriel James Goenawan +4
A recent work shows that Attention Transfer, which transfers only the attention patterns from a pre-trained teacher Vision Transformer (ViT) to a randomly initialized standard stud…
cs.CV2025
Beyond Instance Consistency: Investigating View Diversity in Self-supervised Learning
Huaiyuan Qin, Muli Yang, Siyuan Hu +4
Self-supervised learning (SSL) conventionally relies on the instance consistency paradigm, assuming that different views of the same image can be treated as positive pairs. However…
cs.CV2024
Hybrid Data-Free Knowledge Distillation
Jialiang Tang, Shuo Chen, Chen Gong
Data-free knowledge distillation aims to learn a compact student network from a pre-trained large teacher network without using the original training data of the teacher network. E…