3 papers
cs.CV2026
Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation
Xunyi Zhao, Sihao Lin, Gengze Zhou +5
Instance Goal Navigation (IGN) requires an embodied agent to find a specific object instance among distractors from an under-specified natural-language description. Such ambiguity…
cs.CV2026
VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents
Xunyi Zhao, Gengze Zhou, Qi Wu
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a wide range of vision-language tasks. However, their performance as embodied agents, whic…
cs.CV2025
VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
Sihao Lin, Zerui Li, Xunyi Zhao +10
Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomin…