3 papers
cs.IR2024
Towards Coarse-grained Visual Language Navigation Task Planning Enhanced by Event Knowledge Graph
Zhao Kaichen, Song Yaoxian, Zhao Haiquan +3
Visual language navigation (VLN) is one of the important research in embodied AI. It aims to enable an agent to understand the surrounding environment and complete navigation tasks…
cs.CV2024
Multi-Task Domain Adaptation for Language Grounding with 3D Objects
Penglei Sun, Yaoxian Song, Xinglin Pan +6
The existing works on object-level language grounding with 3D objects mostly focus on improving performance by utilizing the off-the-shelf pre-trained models to capture features, s…
cs.IR2024
Flickr30K-CFQ: A Compact and Fragmented Query Dataset for Text-image Retrieval
Haoyu Liu, Yaoxian Song, Xuwu Wang +4
With the explosive growth of multi-modal information on the Internet, unimodal search cannot satisfy the requirement of Internet applications. Text-image retrieval research is need…