95 citations · 115 across the 4 of their papers we have counts for
1 paper · 1 filter
Youwei Liang, Chongjian Ge, Zhan Tong +3
Vision Transformers (ViTs) take all the image patches as tokens and construct multi-head self-attention (MHSA) among them. Complete leverage of these image tokens brings redundant…