Publications (20)
Vision Transformer with Sparse Scan Prior
Yuguang Zhang, Qihang Fan, Huaibo Huang
In recent years, Transformers have achieved remarkable progress in computer vision tasks. However, their global modeling often comes with substantial computational overhead, in sta…
Lightweight Vision Transformer with Bidirectional Interaction
Qihang Fan, Huaibo Huang, Xiaoqiang Zhou +1
Recent advancements in vision backbones have significantly improved their performance by simultaneously modeling images' local and global contexts. However, the bidirectional inter…
Rethinking Local Perception in Lightweight Vision Transformer
Qihang Fan, Huaibo Huang, Jiyang Guan +1
Vision Transformers (ViTs) have been shown to be effective in various vision tasks. However, resizing them to a mobile-friendly size leads to significant performance degradation. T…
DeVAn: Dense Video Annotation for Video-Language Models
Tingkai Liu, Yunzhe Tao, Haogeng Liu +5
We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeV…
Band-Attention Modulated RetNet for Face Forgery Detection
Zhida Zhang, Jie Cao, Wenkui Yang +3
The transformer networks are extensively utilized in face forgery detection due to their scalability across large datasets.Despite their success, transformers face challenges in ba…
Random Wins All: Rethinking Grouping Strategies for Vision Tokens
Qihang Fan, Yuang Ai, Huaibo Huang +1
Since Transformers are introduced into vision architectures, their quadratic complexity has always been a significant issue that many research efforts aim to address. A representat…