33 citations · 41 across the 6 of their papers we have counts for
8 papers · 1 filter
Prompt-Free Universal Region Proposal Network
Qihong Tang, Changhan Liu, Shaofeng Zhang +3
Identifying potential objects is critical for object recognition and analysis across various computer vision applications. Existing methods typically localize potential objects by…
Stable Segment Anything Model
Qi Fan, Xin Tao, Lei Ke +6
The Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to…
UniBoost: Unsupervised Unimodal Pre-training for Boosting Zero-shot Vision-Language Tasks
Yanan Sun, Zihan Zhong, Qi Fan +2
Large-scale joint training of multimodal models, e.g., CLIP, have demonstrated great performance in many vision-language tasks. However, image-text pairs for pre-training are restr…
Rethinking Local Perception in Lightweight Vision Transformer
Qihang Fan, Huaibo Huang, Jiyang Guan +1
Vision Transformers (ViTs) have been shown to be effective in various vision tasks. However, resizing them to a mobile-friendly size leads to significant performance degradation. T…
Normalization Perturbation: A Simple Domain Generalization Method for Real-World Domain Shifts
Qi Fan, Mattia Segu, Yu-Wing Tai +4
Improving model's generalizability against domain shifts is crucial, especially for safety-critical applications such as autonomous driving. Real-world domain styles can vary subst…
Group Collaborative Learning for Co-Salient Object Detection
Qi Fan, Deng-Ping Fan, Huazhu Fu +3
We present a novel group collaborative learning framework (GCoNet) capable of detecting co-salient objects in real time (16ms), by simultaneously mining consensus representations a…