activity
20242026
collaborators

8 papers

cs.RO2026

Enhancing Vision-Language Navigation with Multimodal Event Knowledge from Real-World Indoor Tour Videos

Haoxuan Xu, Tianfu Li, Wenbo Chen +4

Vision-Language Navigation (VLN) agents often struggle with long-horizon reasoning in unseen environments, particularly when facing ambiguous, coarse-grained instructions. While re…

cs.CV2025

City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning

Penglei Sun, Yaoxian Song, Xiangru Zhu +7

Scene understanding enables intelligent agents to interpret and comprehend their environment. While existing large vision-language models (LVLMs) for scene understanding have prima…

cs.CL2025

Evaluating Semantic Variation in Text-to-Image Synthesis: A Causal Perspective

Xiangru Zhu, Penglei Sun, Yaoxian Song +6

Accurate interpretation and visualization of human instructions are crucial for text-to-image (T2I) synthesis. However, current models struggle to capture semantic variations from…

cs.IR2024

Towards Coarse-grained Visual Language Navigation Task Planning Enhanced by Event Knowledge Graph

Zhao Kaichen, Song Yaoxian, Zhao Haiquan +3

Visual language navigation (VLN) is one of the important research in embodied AI. It aims to enable an agent to understand the surrounding environment and complete navigation tasks…

cs.CV2024

3D Question Answering for City Scene Understanding

Penglei Sun, Yaoxian Song, Xiang Liu +5

3D multimodal question answering (MQA) plays a crucial role in scene understanding by enabling intelligent agents to comprehend their surroundings in 3D environments. While existin…

cs.CV2024

Multi-Task Domain Adaptation for Language Grounding with 3D Objects

Penglei Sun, Yaoxian Song, Xinglin Pan +6

The existing works on object-level language grounding with 3D objects mostly focus on improving performance by utilizing the off-the-shelf pre-trained models to capture features, s…