2 papers
cs.IR2024
Flickr30K-CFQ: A Compact and Fragmented Query Dataset for Text-image Retrieval
Haoyu Liu, Yaoxian Song, Xuwu Wang +4
With the explosive growth of multi-modal information on the Internet, unimodal search cannot satisfy the requirement of Internet applications. Text-image retrieval research is need…
cs.CL2024
OVEL: Large Language Model as Memory Manager for Online Video Entity Linking
Haiquan Zhao, Xuwu Wang, Shisong Chen +3
In recent years, multi-modal entity linking (MEL) has garnered increasing attention in the research community due to its significance in numerous multi-modal applications. Video, a…