From the 1 of 4 linked papers with an AI index.
4 papers
Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation
Jian Zhou, Xunyi Zhao, Gengze Zhou +4
The paper investigates using general-purpose language models as autonomous embodied agents for zero-shot vision-and-language navigation, showing that with only a monocular RGB came…
Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation
Xunyi Zhao, Sihao Lin, Gengze Zhou +5
Instance Goal Navigation (IGN) requires an embodied agent to find a specific object instance among distractors from an under-specified natural-language description. Such ambiguity…
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents
Xunyi Zhao, Gengze Zhou, Qi Wu
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a wide range of vision-language tasks. However, their performance as embodied agents, whic…
VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
Sihao Lin, Zerui Li, Xunyi Zhao +10
Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomin…