collaborators

6 papers

cs.RO2026

MapDream: Task-Driven Map Learning for Vision-Language Navigation

Guoxin Lian, Shuo Wang, Yucheng Wang +7

Vision-Language Navigation (VLN) requires agents to follow natural language instructions in partially observed 3D environments, motivating map representations that aggregate spatia…

cs.RO2026

HoloMotion-1 Technical Report

Maiyue Chen, Kaihui Wang, Bo Zhang +7

In this report, we present HoloMotion-1, a humanoid motion foundation model for zero-shot whole-body motion tracking. A key innovation of HoloMotion-1 is to scale control-policy tr…

cs.RO2026

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation

Shuo Wang, Yucheng Wang, Guoxin Lian +9

Vision-Language Navigation requires agents to act coherently over long horizons by understanding not only local visual context but also how far they have advanced within a multi-st…

cs.RO2026

Scaling Sim-to-Real Reinforcement Learning for Robot VLAs with Generative 3D Worlds

Andrew Choi, Xinjie Wang, Zhizhong Su +1

The strong performance of large vision-language models (VLMs) trained with reinforcement learning (RL) has motivated similar approaches for fine-tuning vision-language-action (VLA)…

cs.CV2025

MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming

Shuo Wang, Yongcai Wang, Zhaoxin Fan +8

Vision-Language Navigation (VLN) tasks often leverage panoramic RGB and depth inputs to provide rich spatial cues for action planning, but these sensors can be costly or less acces…

cs.RO2025

Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation

Shuo Wang, Yongcai Wang, Wanting Li +7

Vision-Language Navigation (VLN) is a critical task for developing embodied agents that can follow natural language instructions to navigate in complex real-world environments. Rec…