activity
20192026
most citedRethinking Local Perception in Lightweight Vision Transformer

33 citations · 41 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

Prompt-Free Universal Region Proposal Network

Qihong Tang, Changhan Liu, Shaofeng Zhang +3

Identifying potential objects is critical for object recognition and analysis across various computer vision applications. Existing methods typically localize potential objects by…

cs.CV2023★ 1 cited

Stable Segment Anything Model

Qi Fan, Xin Tao, Lei Ke +6

The Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to…

cs.CV2023★ 1 cited

UniBoost: Unsupervised Unimodal Pre-training for Boosting Zero-shot Vision-Language Tasks

Yanan Sun, Zihan Zhong, Qi Fan +2

Large-scale joint training of multimodal models, e.g., CLIP, have demonstrated great performance in many vision-language tasks. However, image-text pairs for pre-training are restr…

cs.CV2023★ 33 cited

Rethinking Local Perception in Lightweight Vision Transformer

Qihang Fan, Huaibo Huang, Jiyang Guan +1

Vision Transformers (ViTs) have been shown to be effective in various vision tasks. However, resizing them to a mobile-friendly size leads to significant performance degradation. T…

cs.CV2022★ 5 cited

Normalization Perturbation: A Simple Domain Generalization Method for Real-World Domain Shifts

Qi Fan, Mattia Segu, Yu-Wing Tai +4

Improving model's generalizability against domain shifts is crucial, especially for safety-critical applications such as autonomous driving. Real-world domain styles can vary subst…

cs.CV2021

Group Collaborative Learning for Co-Salient Object Detection

Qi Fan, Deng-Ping Fan, Huazhu Fu +3

We present a novel group collaborative learning framework (GCoNet) capable of detecting co-salient objects in real time (16ms), by simultaneously mining consensus representations a…