activity
20202025
most citedAgent Attention: On the Integration of Softmax and Linear Attention

22 citations · 90 across the 26 of their papers we have counts for

collaborators
Showing cs.CVShow all

26 papers · 1 filter

cs.CV2025

Emulating Human-like Adaptive Vision for Efficient and Flexible Machine Visual Perception

Yulin Wang, Yang Yue, Huanqian Wang +11

Human vision is highly adaptive, efficiently sampling intricate environments by sequentially fixating on task-relevant regions. In contrast, prevailing machine vision models passiv…

cs.CV2024

Bridging the Divide: Reconsidering Softmax and Linear Attention

Dongchen Han, Yifan Pu, Zhuofan Xia +6

Widely adopted in modern Vision Transformer designs, Softmax attention can effectively capture long-range visual information; however, it incurs excessive computational cost when d…

cs.CV2024

A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs

Wangbo Zhao, Yizeng Han, Jiasheng Tang +5

Vision-language models (VLMs) have shown remarkable success across various multi-modal tasks, yet large VLMs encounter significant efficiency challenges due to processing numerous…

cs.CV2024

ENAT: Rethinking Spatial-temporal Interactions in Token-based Image Synthesis

Zanlin Ni, Yulin Wang, Renping Zhou +5

Recently, token-based generation have demonstrated their effectiveness in image synthesis. As a representative example, non-autoregressive Transformers (NATs) can generate decent-q…

cs.CV2024

Exploring contextual modeling with linear complexity for point cloud segmentation

Yong Xien Chng, Xuchong Qiu, Yizeng Han +3

Point cloud segmentation is an important topic in 3D understanding that has traditionally has been tackled using either the CNN or Transformer. Recently, Mamba has emerged as a pro…

cs.CV2024★ 1 cited

Adapting Vision-Language Model with Fine-grained Semantics for Open-Vocabulary Segmentation

Yong Xien Chng, Xuchong Qiu, Yizeng Han +3

Despite extensive research, open-vocabulary segmentation methods still struggle to generalize across diverse domains. To reduce the computational cost of adapting Vision-Language M…