3 papers
cs.RO2026
HCSG: Human-Centric Semantic-Geometric Reasoning for Vision-Language Navigation
Haoxuan Xu, Tianfu Li, Wenbo Chen +7
VLN has achieved remarkable progress by scaling data and model capacity. However, the assumption of a static environment breaks down in real-world indoor scenarios, where robots in…
cs.RO2026
PNav: End-to-End Perception, Prediction and Planning for Vision-and-Language Navigation
Tianfu Li, Wenbo Chen, Haoxuan Xu +2
In Vision-and-Language Navigation (VLN), an agent is required to plan a path to the target specified by the language instruction, using its visual observations. Consequently, preva…
cs.RO2026
Enhancing Vision-Language Navigation with Multimodal Event Knowledge from Real-World Indoor Tour Videos
Haoxuan Xu, Tianfu Li, Wenbo Chen +4
Vision-Language Navigation (VLN) agents often struggle with long-horizon reasoning in unseen environments, particularly when facing ambiguous, coarse-grained instructions. While re…