3 citations · 6 across the 8 of their papers we have counts for
9 papers · 1 filter
Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
Weihao Cao, Runqi Wang, Xiaoyue Duan +3
Open-vocabulary object detection (OVOD) enables models to detect any object category, including unseen ones. Benefiting from large-scale pre-training, existing OVOD methods achieve…
WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
Haodong Zhu, Wenhao Dong, Linlin Yang +10
Leveraging the complementary characteristics of visible (RGB) and infrared (IR) imagery offers significant potential for improving object detection. In this paper, we propose WaveM…
P4Q: Learning to Prompt for Quantization in Visual-language Models
Huixin Sun, Runqi Wang, Yanjing Li +4
Large-scale pre-trained Vision-Language Models (VLMs) have gained prominence in various visual and multimodal tasks, yet the deployment of VLMs on downstream application platforms…
Cross-Level Distillation and Feature Denoising for Cross-Domain Few-Shot Classification
Hao Zheng, Runqi Wang, Jianzhuang Liu +1
The conventional few-shot classification aims at learning a model on a large labeled base dataset and rapidly adapting to a target dataset that is from the same distribution as the…
Self-Enhancement Improves Text-Image Retrieval in Foundation Visual-Language Models
Yuguang Yang, Yiming Wang, Shupeng Geng +4
The emergence of cross-modal foundation models has introduced numerous approaches grounded in text-image retrieval. However, on some domain-specific retrieval tasks, these models f…
Few-Shot Learning with Visual Distribution Calibration and Cross-Modal Distribution Alignment
Runqi Wang, Hao Zheng, Xiaoyue Duan +5
Pre-trained vision-language models have inspired much research on few-shot learning. However, with only a few training images, there exist two crucial problems: (1) the visual feat…