22 citations · 91 across the 27 of their papers we have counts for
12 papers · 2 filters
Bridging the Divide: Reconsidering Softmax and Linear Attention
Dongchen Han, Yifan Pu, Zhuofan Xia +6
Widely adopted in modern Vision Transformer designs, Softmax attention can effectively capture long-range visual information; however, it incurs excessive computational cost when d…
A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs
Wangbo Zhao, Yizeng Han, Jiasheng Tang +5
Vision-language models (VLMs) have shown remarkable success across various multi-modal tasks, yet large VLMs encounter significant efficiency challenges due to processing numerous…
ENAT: Rethinking Spatial-temporal Interactions in Token-based Image Synthesis
Zanlin Ni, Yulin Wang, Renping Zhou +5
Recently, token-based generation have demonstrated their effectiveness in image synthesis. As a representative example, non-autoregressive Transformers (NATs) can generate decent-q…
Exploring contextual modeling with linear complexity for point cloud segmentation
Yong Xien Chng, Xuchong Qiu, Yizeng Han +3
Point cloud segmentation is an important topic in 3D understanding that has traditionally has been tackled using either the CNN or Transformer. Recently, Mamba has emerged as a pro…
Adapting Vision-Language Model with Fine-grained Semantics for Open-Vocabulary Segmentation
Yong Xien Chng, Xuchong Qiu, Yizeng Han +3
Despite extensive research, open-vocabulary segmentation methods still struggle to generalize across diverse domains. To reduce the computational cost of adapting Vision-Language M…
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
Yifan Pu, Zhuofan Xia, Jiayi Guo +9
This paper identifies significant redundancy in the query-key interactions within self-attention mechanisms of diffusion transformer models, particularly during the early stages of…