1 paper
Xiangyun Huang, Xiangchen Wang, Runfeng Lin +5
Vision-and-Language Navigation (VLN) requires an agent to follow a route-level instruction by executing its constituent steps from egocentric visual observations. Existing VLM-base…