10 papers · 1 filter
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
Xinghang Li, Jun Guo, Qiwei Li +21
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements f…
Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation
Nan Sun, Yuan Zhang, Yongkun Yang +10
Embodied chain-of-thought (CoT) aims to bridge linguistic reasoning and robotic control, but its effective form and integration strategy remain underexplored. In this paper, we rev…
RotVLA: Rotational Latent Action for Vision-Language-Action Model
Qiwei Li, Xicheng Gong, Xinghang Li +5
Latent Action Models (LAMs) have emerged as an effective paradigm for handling heterogeneous datasets during Vision-Language-Action (VLA) model pretraining, offering a unified acti…
Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
Jun Guo, Qiwei Li, Peiyan Li +7
We propose X-WAM, a Unified 4D World Model that unifies real-time robotic action execution and high-fidelity 4D world synthesis (video + 3D reconstruction) in a single framework, a…
SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy
Peiyan Li, Yixiang Chen, Yuan Xu +13
Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies neglect one or both aspects. The…
Scaling World Model for Hierarchical Manipulation Policies
Qian Long, Yueze Wang, Jiaxi Song +13
Vision-Language-Action (VLA) models are promising for generalist robot manipulation but remain brittle in out-of-distribution (OOD) settings, especially with limited real-robot dat…