works on

From the 1 of 8 linked papers with an AI index.

collaborators

8 papers

cs.RO2026

Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation

Jian Zhou, Xunyi Zhao, Gengze Zhou +4

The paper investigates using general-purpose language models as autonomous embodied agents for zero-shot vision-and-language navigation, showing that with only a monocular RGB came…

cs.RO2026

Automating the Design of Embodied Agent Architectures

Jian Zhou, Sihao Lin, Jin Li +3

Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules. This modularity exposes a large architectural design space, b…

cs.RO2026

One Agent to Guide Them All: Empowering MLLMs for Vision-and-Language Navigation via Explicit World Representation

Zerui Li, Hongpei Zheng, Fangguo Zhao +5

A navigable agent needs to understand both high-level semantic instructions and precise spatial perceptions. Building navigation agents centered on Multimodal Large Language Models…

cs.CV2026

VLN-MME: Diagnosing MLLMs as Language-guided Visual Navigation agents

Xunyi Zhao, Gengze Zhou, Qi Wu

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a wide range of vision-language tasks. However, their performance as embodied agents, whic…

cs.CV2025

VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation

Sihao Lin, Zerui Li, Xunyi Zhao +10

Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomin…

cs.CL2025

Ask Good Questions for Large Language Models

Qi Wu, Zhongqi Lu

Recent advances in large language models (LLMs) have significantly improved the performance of dialog systems, yet current approaches often fail to provide accurate guidance of top…