1 paper
Zhe Bian, Zhe Wang, Wenqiang Han +1
Since its inception, Vision Transformer (ViT) has emerged as a prevalent model in the computer vision domain. Nonetheless, the multi-head self-attention (MHSA) mechanism in ViT is…