activity
20242026
collaborators

12 papers

cs.RO2026

From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation

Xiangyu Shi, Ruoxi Yang, Wei Tao +3

Vision-and-Language Navigation (VLN) agents may satisfy conventional success criteria while still failing to establish reliable object-level grounding, because current evaluation p…

cs.RO2025

Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation

Xiangyu Shi, Zerui Li, Yanyuan Qiao +1

Recent advances in Vision-and-Language Navigation in Continuous Environments (VLN-CE) have leveraged multimodal large language models (MLLMs) to achieve zero-shot navigation. Howev…

cs.AI2025

Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs

Yue Zhang, Tianyi Ma, Zun Wang +2

Integrating large language models (LLMs) into embodied AI models is becoming increasingly prevalent. However, existing zero-shot LLM-based Vision-and-Language Navigation (VLN) agen…

cs.RO2025

SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation

Xiangyu Shi, Zerui Li, Wenqi Lyu +4

Vision-and-Language Navigation (VLN) in continuous environments requires agents to interpret natural language instructions while navigating unconstrained 3D spaces. Existing VLN-CE…

cs.CV2025

NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Yanyuan Qiao, Haodong Hong, Wenqi Lyu +5

Multimodal Large Language Models (MLLMs) have demonstrated strong generalization in vision-language tasks, yet their ability to understand and act within embodied environments rema…

cs.RO2025

BadNAVer: Exploring Jailbreak Attacks On Vision-and-Language Navigation

Wenqi Lyu, Zerui Li, Yanyuan Qiao +1

Multimodal large language models (MLLMs) have recently gained attention for their generalization and reasoning capabilities in Vision-and-Language Navigation (VLN) tasks, leading t…