3 papers
cs.IR2024
Flickr30K-CFQ: A Compact and Fragmented Query Dataset for Text-image Retrieval
Haoyu Liu, Yaoxian Song, Xuwu Wang +4
With the explosive growth of multi-modal information on the Internet, unimodal search cannot satisfy the requirement of Internet applications. Text-image retrieval research is need…
cs.AI2023
M^2ConceptBase: A Fine-Grained Aligned Concept-Centric Multimodal Knowledge Base
Zhiwei Zha, Jiaan Wang, Zhixu Li +3
Multimodal knowledge bases (MMKBs) provide cross-modal aligned knowledge crucial for multimodal tasks. However, the images in existing MMKBs are generally collected for entities in…
cs.AI2023
Scene-Driven Multimodal Knowledge Graph Construction for Embodied AI
Song Yaoxian, Sun Penglei, Liu Haoyu +4
Embodied AI is one of the most popular studies in artificial intelligence and robotics, which can effectively improve the intelligence of real-world agents (i.e. robots) serving hu…