28 papers
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Xiaomi Robotics Team, Jun Guo, Piaopiao Jin +31
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulatio…
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
Xinghang Li, Jun Guo, Qiwei Li +21
The paper introduces Xiaomi-Robotics-U0, a 38‑billion‑parameter multimodal autoregressive model that extends foundation image and video generation to embodied robotics, enabling co…
BabyVision: Visual Reasoning Beyond Language
Liang Chen, Weichu Xie, Yiyan Liang +27
While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile…
RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark
Huashuo Lei, Wenxuan Song, Huarui Zhang +10
Memory is a critical component of robotic intelligence, as robots must rely on past observations and actions to accomplish long-horizon tasks in partially observable environments.…
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
Wenxuan Song, Han Zhao, Fuhao Li +7
This paper proposes a novel approach to address the challenge that pretrained VLA models often fail to effectively improve performance and reduce adaptation costs during standard s…
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
Jiayi Chen, Wenxuan Song, Pengxiang Ding +5
Vision-language-action (VLA) models aim to understand natural language instructions and visual observations and to execute corresponding actions as an embodied agent. Recent work i…