4 papers
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
Penglei Sun, Yaoxian Song, Xiangru Zhu +7
Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have prima…
Towards Coarse-grained Visual Language Navigation Task Planning Enhanced by Event Knowledge Graph
Zhao Kaichen, Song Yaoxian, Zhao Haiquan +3
Visual language navigation (VLN) is one of the important research in embodied AI. It aims to enable an agent to understand the surrounding environment and complete navigation tasks…
3D Question Answering for City Scene Understanding
Penglei Sun, Yaoxian Song, Xiang Liu +5
3D multimodal question answering (MQA) plays a crucial role in scene understanding by enabling intelligent agents to comprehend their surroundings in 3D environments. While existin…
Multi-Task Domain Adaptation for Language Grounding with 3D Objects
Penglei Sun, Yaoxian Song, Xinglin Pan +6
The existing works on object-level language grounding with 3D objects mostly focus on improving performance by utilizing the off-the-shelf pre-trained models to capture features, s…