18 papers · 1 filter
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Xiaomi Robotics Team, Jun Guo, Piaopiao Jin +31
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulatio…
FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation
Lingfeng Zhang, Zeying Gong, Xiaoshuai Hao +7
Vision-and-language navigation (VLN) in continuous environments requires an agent to ground instructions in egocentric observations while maintaining spatial understanding across l…
Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning
Ziyang Yao, Haochen Liu, Yuncheng Jiang +10
Autonomous driving requires reasoning about how ego actions shape future world evolution, rather than merely mapping observations to actions. However, most end-to-end methods rely…
OneVLA: A Unified Framework for Embodied Tasks
Lingfeng Zhang, Xiaoshuai Hao, Yingbo Tang +10
Navigation and manipulation are fundamental capabilities of embodied intelligence, enabling robots to interpret natural language commands and interact physically with their surroun…
SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework
Tianshu Wu, Xiangqi Kong, Yue Chen +5
Building humanoid robots capable of generalizable whole-body loco-manipulation in the real world remains a fundamental challenge. Existing methods either rely on laborious task-spe…
Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives
Junli Wang, Zhihua Hua, Xueyi Liu +7
Existing imitation learning methods for end-to-end autonomous driving predominantly learn from successful demonstrations by minimizing geometric deviations from expert trajectories…