22 citations · 90 across the 26 of their papers we have counts for
26 papers · 1 filter
Emulating Human-like Adaptive Vision for Efficient and Flexible Machine Visual Perception
Yulin Wang, Yang Yue, Huanqian Wang +11
Human vision is highly adaptive, efficiently sampling intricate environments by sequentially fixating on task-relevant regions. In contrast, prevailing machine vision models passiv…
Bridging the Divide: Reconsidering Softmax and Linear Attention
Dongchen Han, Yifan Pu, Zhuofan Xia +6
Widely adopted in modern Vision Transformer designs, Softmax attention can effectively capture long-range visual information; however, it incurs excessive computational cost when d…
A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs
Wangbo Zhao, Yizeng Han, Jiasheng Tang +5
Vision-language models (VLMs) have shown remarkable success across various multi-modal tasks, yet large VLMs encounter significant efficiency challenges due to processing numerous…
ENAT: Rethinking Spatial-temporal Interactions in Token-based Image Synthesis
Zanlin Ni, Yulin Wang, Renping Zhou +5
Recently, token-based generation have demonstrated their effectiveness in image synthesis. As a representative example, non-autoregressive Transformers (NATs) can generate decent-q…
Exploring contextual modeling with linear complexity for point cloud segmentation
Yong Xien Chng, Xuchong Qiu, Yizeng Han +3
Point cloud segmentation is an important topic in 3D understanding that has traditionally has been tackled using either the CNN or Transformer. Recently, Mamba has emerged as a pro…
Adapting Vision-Language Model with Fine-grained Semantics for Open-Vocabulary Segmentation
Yong Xien Chng, Xuchong Qiu, Yizeng Han +3
Despite extensive research, open-vocabulary segmentation methods still struggle to generalize across diverse domains. To reduce the computational cost of adapting Vision-Language M…