3 citations · 3 across the 6 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
HSMLA: Hierarchical Softmax Multi-scale Linear Attention for Efficient Vision Transformers
Dong Liu, Yanxuan Yu, Renata Borovica-Gajic +2
Vision transformers face significant computational overheads in high-resolution dense prediction due to the quadratic complexity of self-attention. Linear attention offers efficien…
cs.CV2024★ 3 cited
Visual Fourier Prompt Tuning
Runjia Zeng, Cheng Han, Qifan Wang +5
With the scale of vision Transformer-based models continuing to grow, finetuning these large-scale pretrained models for new tasks has become increasingly parameter-intensive. Visu…
cs.CV2024★ 10 cited
Image Translation as Diffusion Visual Programmers
Cheng Han, James C. Liang, Qifan Wang +5
We introduce the novel Diffusion Visual Programmer (DVP), a neuro-symbolic image translation framework. Our proposed DVP seamlessly embeds a condition-flexible diffusion model with…