1 paper
Xu Pan, Aaron Philip, Ziqian Xie +1
Self-attention in vision transformers is often thought to perform perceptual grouping where tokens attend to other tokens with similar embeddings, which could correspond to semanti…