3 papers
cs.CV2025
MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
Minghui Hou, Wei-Hsing Huang, Shaofeng Liang +5
Vision-language models enable the understanding and reasoning of complex traffic scenarios through multi-source information fusion, establishing it as a core technology for autonom…
cs.IR2025
CART: A Generative Cross-Modal Retrieval Framework with Coarse-To-Fine Semantic Modeling
Minghui Fang, Shengpeng Ji, Jialong Zuo +9
Cross-modal retrieval aims to search for instances, which are semantically related to the query through the interaction of different modal data. Traditional solutions utilize a sin…
cs.CV2025
EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models
GuangHao Meng, Sunan He, Jinpeng Wang +7
Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries to retrieve corresponding images…