1 paper
Nam Hyeon-Woo, Kim Yu-Ji, Byeongho Heo +3
The favorable performance of Vision Transformers (ViTs) is often attributed to the multi-head self-attention (MSA). The MSA enables global interactions at each layer of a ViT model…