18 citations · 20 across the 4 of their papers we have counts for
4 papers
Multimodal Large Language Model for Visual Navigation
Yao-Hung Hubert Tsai, Vansh Dhar, Jialu Li +2
Recent efforts to enable visual navigation using large language models have mainly focused on developing complex prompt systems. These systems incorporate instructions, observation…
PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation
Jialu Li, Mohit Bansal
Vision-and-Language Navigation (VLN) requires the agent to follow language instructions to navigate through 3D environments. One main challenge in VLN is the limited availability o…
Improving Vision-and-Language Navigation by Generating Future-View Image Semantics
Jialu Li, Mohit Bansal
Vision-and-Language Navigation (VLN) is the task that requires an agent to navigate through the environment based on natural language instructions. At each step, the agent takes th…
CLEAR: Improving Vision-Language Navigation with Cross-Lingual, Environment-Agnostic Representations
Jialu Li, Hao Tan, Mohit Bansal
Vision-and-Language Navigation (VLN) tasks require an agent to navigate through the environment based on language instructions. In this paper, we aim to solve two key challenges in…