3 citations · 4 across the 5 of their papers we have counts for
5 papers · 1 filter
GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
Jiahao Yang, Zihan Wang, Xiangyang Li +4
Despite significant progress in Vision-Language Navigation (VLN), existing approaches still rely on dense RGB videos that produce excessive patch tokens and lack explicit spatial s…
SVLL: Staged Vision-Language Learning for Physically Grounded Embodied Task Planning
Yuyuan Yang, Junkun Hong, Hongrong Wang +11
Embodied task planning demands vision-language models to generate action sequences that are both visually grounded and causally coherent over time. However, existing training parad…
Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language Navigation
Zihan Wang, Xiangyang Li, Jiahao Yang +4
Vision-and-language navigation (VLN) enables the agent to navigate to a remote location following the natural language instruction in 3D environments. At each navigation step, the…
GridMM: Grid Memory Map for Vision-and-Language Navigation
Zihan Wang, Xiangyang Li, Jiahao Yang +2
Vision-and-language navigation (VLN) enables the agent to navigate to a remote location following the natural language instruction in 3D environments. To represent the previously v…
KERM: Knowledge Enhanced Reasoning for Vision-and-Language Navigation
Xiangyang Li, Zihan Wang, Jiahao Yang +2
Vision-and-language navigation (VLN) is the task to enable an embodied agent to navigate to a remote location following the natural language instruction in real scenes. Most of the…