activity
20242026
collaborators

7 papers

cs.CV2026

VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents

Xunyi Zhao, Gengze Zhou, Qi Wu

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a wide range of vision-language tasks. However, their performance as embodied agents, whic…

cs.CV2025

VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation

Sihao Lin, Zerui Li, Xunyi Zhao +10

Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomin…

cs.CL2025

MMGR: Multi-Modal Generative Reasoning

Zefan Cai, Haoyi Qiu, Tianyi Ma +9

Video foundation models generate visually realistic and temporally coherent content, but their reliability as world simulators depends on whether they capture physical, logical, an…

cs.CV2025

Rethinking Training Dynamics in Scale-wise Autoregressive Generation

Gengze Zhou, Chongjian Ge, Hao Tan +2

Recent advances in autoregressive (AR) generative models have produced increasingly powerful systems for media synthesis. Among them, next-scale prediction has emerged as a popular…

cs.RO2025

Embodied Navigation Foundation Model

Jiazhao Zhang, Anqi Li, Yunpeng Qi +14

Navigation is a fundamental capability in embodied AI, representing the intelligence required to perceive and interact within physical environments following language instructions.…

cs.RO2025

Ground-level Viewpoint Vision-and-Language Navigation in Continuous Environments

Zerui Li, Gengze Zhou, Haodong Hong +4

Vision-and-Language Navigation (VLN) empowers agents to associate time-sequenced visual observations with corresponding instructions to make sequential decisions. However, generali…