4 papers
The Solution for CVPR2024 Foundational Few-Shot Object Detection Challenge
Hongpeng Pan, Shifeng Yi, Shouwei Yang +4
This report introduces an enhanced method for the Foundational Few-Shot Object Detection (FSOD) task, leveraging the vision-language model (VLM) for object detection. However, on s…
Robust Semi-supervised Learning by Wisely Leveraging Open-set Data
Yang Yang, Nan Jiang, Yi Xu +1
Open-set Semi-supervised Learning (OSSL) holds a realistic setting that unlabeled data may come from classes unseen in the labeled set, i.e., out-of-distribution (OOD) data, which…
Learning to Rebalance Multi-Modal Optimization by Adaptively Masking Subnetworks
Yang Yang, Hongpeng Pan, Qing-Yuan Jiang +2
Multi-modal learning aims to enhance performance by unifying models from various modalities but often faces the "modality imbalance" problem in real data, leading to a bias towards…
The Solution for the CVPR 2023 1st foundation model challenge-Track2
Haonan Xu, Yurui Huang, Sishun Pan +3
In this paper, we propose a solution for cross-modal transportation retrieval. Due to the cross-domain problem of traffic images, we divide the problem into two sub-tasks of pedest…