Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
The Solution for CVPR2024 Foundational Few-Shot Object Detection Challenge
Hongpeng Pan, Shifeng Yi, Shouwei Yang +4
This report introduces an enhanced method for the Foundational Few-Shot Object Detection (FSOD) task, leveraging the vision-language model (VLM) for object detection. However, on s…
cs.CV2024
Learning to Rebalance Multi-Modal Optimization by Adaptively Masking Subnetworks
Yang Yang, Hongpeng Pan, Qing-Yuan Jiang +2
Multi-modal learning aims to enhance performance by unifying models from various modalities but often faces the "modality imbalance" problem in real data, leading to a bias towards…