collaborators

7 papers

cs.CV2026

MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

Zehua Fan, Junjie He, Wenxuan Song +14

World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands sim…

cs.CV2026

PROSPECT: Unified Streaming Vision-Language Navigation via Semantic--Spatial Fusion and Latent Predictive Representation

Zehua Fan, Wenqi Lyu, Wenxuan Song +12

Multimodal large language models (MLLMs) have advanced zero-shot end-to-end Vision-Language Navigation (VLN), yet robust navigation requires not only semantic understanding but als…

cs.RO2025

SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation

Xiangyu Shi, Zerui Li, Wenqi Lyu +4

Vision-and-Language Navigation (VLN) in continuous environments requires agents to interpret natural language instructions while navigating unconstrained 3D spaces. Existing VLN-CE…

cs.CV2025

NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Yanyuan Qiao, Haodong Hong, Wenqi Lyu +5

Multimodal Large Language Models (MLLMs) have demonstrated strong generalization in vision-language tasks, yet their ability to understand and act within embodied environments rema…

cs.RO2025

BadNAVer: Exploring Jailbreak Attacks On Vision-and-Language Navigation

Wenqi Lyu, Zerui Li, Yanyuan Qiao +1

Multimodal large language models (MLLMs) have recently gained attention for their generalization and reasoning capabilities in Vision-and-Language Navigation (VLN) tasks, leading t…

cs.RO2025

Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments

Zerui Li, Gengze Zhou, Haodong Hong +4

Vision-and-Language Navigation (VLN) empowers agents to associate time-sequenced visual observations with corresponding instructions to make sequential decisions. However, generali…