3 papers
cs.CV2026
GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
Jiahao Yang, Zihan Wang, Xiangyang Li +4
Despite significant progress in Vision-Language Navigation (VLN), existing approaches still rely on dense RGB videos that produce excessive patch tokens and lack explicit spatial s…
cs.CV2026
SVLL: Staged Vision-Language Learning for Physically Grounded Embodied Task Planning
Yuyuan Yang, Junkun Hong, Hongrong Wang +11
Embodied task planning demands vision-language models to generate action sequences that are both visually grounded and causally coherent over time. However, existing training parad…
cs.RO2024
Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation
Zihan Wang, Xiangyang Li, Jiahao Yang +2
Vision-and-language navigation (VLN) enables the agent to navigate to a remote location in 3D environments following the natural language instruction. In this field, the agent is u…