papers

Publications (20)

cs.CV2025

Vision Transformer with Sparse Scan Prior

Yuguang Zhang, Qihang Fan, Huaibo Huang

In recent years, Transformers have achieved remarkable progress in computer vision tasks. However, their global modeling often comes with substantial computational overhead, in sta…

cs.CV2025

Lightweight Vision Transformer with Bidirectional Interaction

Qihang Fan, Huaibo Huang, Xiaoqiang Zhou +1

Recent advancements in vision backbones have significantly improved their performance by simultaneously modeling images' local and global contexts. However, the bidirectional inter…

cs.CV2023

Rethinking Local Perception in Lightweight Vision Transformer

Qihang Fan, Huaibo Huang, Jiyang Guan +1

Vision Transformers (ViTs) have been shown to be effective in various vision tasks. However, resizing them to a mobile-friendly size leads to significant performance degradation. T…

cs.CV2024

DeVAn: Dense Video Annotation for Video-Language Models

Tingkai Liu, Yunzhe Tao, Haogeng Liu +5

We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeV…

cs.CV2024

Band-Attention Modulated RetNet for Face Forgery Detection

Zhida Zhang, Jie Cao, Wenkui Yang +3

The transformer networks are extensively utilized in face forgery detection due to their scalability across large datasets.Despite their success, transformers face challenges in ba…

cs.CV2026

Random Wins All: Rethinking Grouping Strategies for Vision Tokens

Qihang Fan, Yuang Ai, Huaibo Huang +1

Since Transformers are introduced into vision architectures, their quadratic complexity has always been a significant issue that many research efforts aim to address. A representat…