8 papers
Enhancing Vision-Language Navigation with Multimodal Event Knowledge from Real-World Indoor Tour Videos
Haoxuan Xu, Tianfu Li, Wenbo Chen +4
Vision-Language Navigation (VLN) agents often struggle with long-horizon reasoning in unseen environments, particularly when facing ambiguous, coarse-grained instructions. While re…
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
Penglei Sun, Yaoxian Song, Xiangru Zhu +7
Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have prima…
Evaluating Semantic Variation in Text-to-Image Synthesis: A Causal Perspective
Xiangru Zhu, Penglei Sun, Yaoxian Song +6
Accurate interpretation and visualization of human instructions are crucial for text-to-image (T2I) synthesis. However, current models struggle to capture semantic variations from…
Towards Coarse-grained Visual Language Navigation Task Planning Enhanced by Event Knowledge Graph
Zhao Kaichen, Song Yaoxian, Zhao Haiquan +3
Visual language navigation (VLN) is one of the important research in embodied AI. It aims to enable an agent to understand the surrounding environment and complete navigation tasks…
3D Question Answering for City Scene Understanding
Penglei Sun, Yaoxian Song, Xiang Liu +5
3D multimodal question answering (MQA) plays a crucial role in scene understanding by enabling intelligent agents to comprehend their surroundings in 3D environments. While existin…
Multi-Task Domain Adaptation for Language Grounding with 3D Objects
Penglei Sun, Yaoxian Song, Xinglin Pan +6
The existing works on object-level language grounding with 3D objects mostly focus on improving performance by utilizing the off-the-shelf pre-trained models to capture features, s…