3 papers
cs.CV2024
Rethinking Misalignment in Vision-Language Model Adaptation from a Causal Perspective
Yanan Zhang, Jiangmeng Li, Lixiang Liu +1
Foundational Vision-Language models such as CLIP have exhibited impressive generalization in downstream tasks. However, CLIP suffers from a two-level misalignment issue, i.e., task…
cs.MM2024
Start from Video-Music Retrieval: An Inter-Intra Modal Loss for Cross Modal Retrieval
Zeyu Chen, Pengfei Zhang, Kai Ye +3
The burgeoning short video industry has accelerated the advancement of video-music retrieval technology, assisting content creators in selecting appropriate music for their videos.…
cs.CL2024
Event-enhanced Retrieval in Real-time Search
Yanan Zhang, Xiaoling Bai, Tianhua Zhou
The embedding-based retrieval (EBR) approach is widely used in mainstream search engine retrieval systems and is crucial in recent retrieval-augmented methods for eliminating LLM i…