5 papers
FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation
Lingfeng Zhang, Zeying Gong, Xiaoshuai Hao +7
Vision-and-language navigation (VLN) in continuous environments requires an agent to ground instructions in egocentric observations while maintaining spatial understanding across l…
VAIC: Vision-Guided Humanoid Agile Object Interaction Control via Decoupled Commands
Dongting Li, Qianyang Wu, Xingyu Chen +9
Humanoid robots hold immense potential for real-world assistance, yet agile interaction with objects in unstructured environments demands tightly coupled whole-body coordination. D…
Learning Diverse Skills for Behavior Models with Mixture of Experts
Wangtian Shen, Jinming Ma, Mingliang Zhou +1
Imitation learning has demonstrated strong performance in robotic manipulation by learning from large-scale human demonstrations. While existing models excel at single-task learnin…
An Efficient and Multi-Modal Navigation System with One-Step World Model
Wangtian Shen, Ziyang Meng, Jinming Ma +2
Navigation is a fundamental capability for mobile robots. While the current trend is to use learning-based approaches to replace traditional geometry-based methods, existing end-to…
Reinforced Embodied Planning with Verifiable Reward for Real-World Robotic Manipulation
Zitong Bo, Yue Hu, Jinming Ma +7
Enabling robots to execute long-horizon manipulation tasks from free-form language instructions remains a fundamental challenge in embodied AI. While vision-language models (VLMs)…