5 papers
From World Models to World Action Models: A Concise Tutorial for Robotics
Xiaoxiong Zhang, Xiong Zeng, Wei Zhang
Rather than providing an exhaustive survey, this paper presents a concise tutorial on world models and world action models for robotics. After reading the tutorial, readers should…
FSUNav: A Cerebrum-Cerebellum Architecture for Fast, Safe, and Universal Zero-Shot Goal-Oriented Navigation
Mingao Tan, Yiyang Li, Shanze Wang +2
Current vision-language navigation methods face substantial bottlenecks regarding heterogeneous robot compatibility, real-time performance, and navigation safety. Furthermore, they…
ECHO: Edge-Cloud Humanoid Orchestration for Language-to-Motion Control
Haozhe Jia, Jianfei Song, Yuan Zhang +5
We present ECHO, an edge--cloud framework for language-driven whole-body control of humanoid robots. A cloud-hosted diffusion-based text-to-motion generator synthesizes motion refe…
Skill-Aware Diffusion for Generalizable Robotic Manipulation
Aoshen Huang, Jiaming Chen, Jiyu Cheng +3
Robust generalization in robotic manipulation is crucial for robots to adapt flexibly to diverse environments. Existing methods usually improve generalization by scaling data and n…
Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation
Chuye Zhang, Xiaoxiong Zhang, Wei Pan +2
Robotic manipulation in unstructured environments requires systems that can generalize across diverse tasks while maintaining robust and reliable performance. We introduce {GVF-TAP…