3 papers
cs.CV2026
DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
Xinwei He, Yansong Zheng, Qianru Han +7
Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging semantically aligned latent…
eess.IV2025
Tetrahedron-Net for Medical Image Registration
Jinhai Xiang, Shuai Guo, Qianru Han +3
Medical image registration plays a vital role in medical image processing. Extracting expressive representations for medical images is crucial for improving the registration qualit…
cs.CV2024
CLIP-SCGI: Synthesized Caption-Guided Inversion for Person Re-Identification
Qianru Han, Xinwei He, Zhi Liu +3
Person re-identification (ReID) has recently benefited from large pretrained vision-language models such as Contrastive Language-Image Pre-Training (CLIP). However, the absence of…