1 paper
Xueying Li, Feng Lyu, Hao Wu +3
Training-free Vision-Language Navigation (VLN) agents powered by foundation models can follow instructions and explore 3D environments. However, existing approaches rely on greedy…