20 citations · 33 across the 4 of their papers we have counts for
4 papers
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
Yifan Pu, Zhuofan Xia, Jiayi Guo +9
This paper identifies significant redundancy in the query-key interactions within self-attention mechanisms of diffusion transformer models, particularly during the early stages of…
DAT++: Spatially Dynamic Vision Transformer with Deformable Attention
Zhuofan Xia, Xuran Pan, Shiji Song +2
Transformers have shown superior performance on various vision tasks. Their large receptive field endows Transformer models with higher representation power than their CNN counterp…
Slide-Transformer: Hierarchical Vision Transformer with Local Self-Attention
Xuran Pan, Tianzhu Ye, Zhuofan Xia +2
Self-attention mechanism has been a key factor in the recent progress of Vision Transformer (ViT), which enables adaptive feature extraction from global contexts. However, existing…
Vision Transformer with Deformable Attention
Zhuofan Xia, Xuran Pan, Shiji Song +2
Transformers have recently shown superior performances on various vision tasks. The large, sometimes even global, receptive field endows Transformer models with higher representati…