Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Linearizing Vision Transformer with Test-Time Training
Yining Li, Dongchen Han, Zeyu Liu +3
While linear-complexity attention mechanisms offer a promising alternative to Softmax attention for overcoming the quadratic bottleneck, training such models from scratch remains p…
cs.CV2026
Linear-Time Global Visual Modeling without Explicit Attention
Ruize He, Dongchen Han, Gao Huang
Existing research largely attributes the global sequence modeling capability of Transformers to the explicit computation of attention weights, a process that inherently incurs quad…
cs.CV2024
VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models
Jiawei Liang, Siyuan Liang, Man Luo +4
Autoregressive Visual Language Models (VLMs) showcase impressive few-shot learning capabilities in a multimodal context. Recently, multimodal instruction tuning has been proposed t…