2 papers
cs.CV2024
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
Dan Song, Xinwei Fu, Ning Liu +5
Large-scale pre-trained models have demonstrated impressive performance in vision and language tasks within open-world scenarios. Due to the lack of comparable pre-trained models f…
cs.CV2024
Towards Deconfounded Image-Text Matching with Causal Inference
Wenhui Li, Xinqi Su, Dan Song +3
Prior image-text matching methods have shown remarkable performance on many benchmark datasets, but most of them overlook the bias in the dataset, which exists in intra-modal and i…