33 citations · 37 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 3 cited
ViTAR: Vision Transformer with Any Resolution
Qihang Fan, Quanzeng You, Xiaotian Han +5
This paper tackles a significant challenge faced by Vision Transformers (ViTs): their constrained scalability across different image resolutions. Typically, ViTs experience a perfo…
cs.CV2023★ 1 cited
Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling
Haogeng Liu, Qihang Fan, Tingkai Liu +5
This paper proposes Video-Teller, a video-language foundation model that leverages multi-modal fusion and fine-grained modality alignment to significantly enhance the video-to-text…
cs.CV2023★ 33 cited
Rethinking Local Perception in Lightweight Vision Transformer
Qihang Fan, Huaibo Huang, Jiyang Guan +1
Vision Transformers (ViTs) have been shown to be effective in various vision tasks. However, resizing them to a mobile-friendly size leads to significant performance degradation. T…