5 papers
DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
Xinwei He, Yansong Zheng, Qianru Han +7
Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging semantically aligned latent…
FF-PNet: A Pyramid Network Based on Feature and Field for Brain Image Registration
Ying Zhang, Shuai Guo, Chenxi Sun +2
In recent years, deformable medical image registration techniques have made significant progress. However, existing models still lack efficiency in parallel extraction of coarse an…
Tetrahedron-Net for Medical Image Registration
Jinhai Xiang, Shuai Guo, Qianru Han +3
Medical image registration plays a vital role in medical image processing. Extracting expressive representations for medical images is crucial for improving the registration qualit…
TeDA: Boosting Vision-Lanuage Models for Zero-Shot 3D Object Retrieval via Testing-time Distribution Alignment
Zhichuan Wang, Yang Zhou, Jinhai Xiang +2
Learning discriminative 3D representations that generalize well to unknown testing categories is an emerging requirement for many real-world 3D applications. Existing well-establis…
CLIP-SCGI: Synthesized Caption-Guided Inversion for Person Re-Identification
Qianru Han, Xinwei He, Zhi Liu +3
Person re-identification (ReID) has recently benefited from large pretrained vision-language models such as Contrastive Language-Image Pre-Training (CLIP). However, the absence of…