20 citations · 49 across the 6 of their papers we have counts for
6 papers
Bridging the Divide: Reconsidering Softmax and Linear Attention
Dongchen Han, Yifan Pu, Zhuofan Xia +6
Widely adopted in modern Vision Transformer designs, Softmax attention can effectively capture long-range visual information; however, it incurs excessive computational cost when d…
DAT++: Spatially Dynamic Vision Transformer with Deformable Attention
Zhuofan Xia, Xuran Pan, Shiji Song +2
Transformers have shown superior performance on various vision tasks. Their large receptive field endows Transformer models with higher representation power than their CNN counterp…
FLatten Transformer: Vision Transformer using Focused Linear Attention
Dongchen Han, Xuran Pan, Yizeng Han +2
The quadratic computation complexity of self-attention has been a persistent challenge when applying Transformer models to vision tasks. Linear attention, on the other hand, offers…
Slide-Transformer: Hierarchical Vision Transformer with Local Self-Attention
Xuran Pan, Tianzhu Ye, Zhuofan Xia +2
Self-attention mechanism has been a key factor in the recent progress of Vision Transformer (ViT), which enables adaptive feature extraction from global contexts. However, existing…
Joint Representation Learning for Text and 3D Point Cloud
Rui Huang, Xuran Pan, Henry Zheng +4
Recent advancements in vision-language pre-training (e.g. CLIP) have shown that vision models can benefit from language supervision. While many models using language modality have…
Vision Transformer with Deformable Attention
Zhuofan Xia, Xuran Pan, Shiji Song +2
Transformers have recently shown superior performances on various vision tasks. The large, sometimes even global, receptive field endows Transformer models with higher representati…