11 papers
CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
Yushan Liu, Peibo Sun, Xintao Chao +8
The paper introduces CheckVLA, a system that uses a frozen action‑conditioned world model to verify and intervene during long‑horizon mobile manipulation when execution deviates fr…
PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving
Yushan Liu, Tianxiong Lv, Bohua Wang +11
Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-relevant cues, while static fusion…
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Xiaomi Robotics Team, Jun Guo, Piaopiao Jin +31
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulatio…
TaCauchy: An Extensible FEM Framework for Vision-Based Tactile Simulation
Hengfei Zhao, Yifan Xie, Junhao Gong +6
Vision-based tactile sensors require high-fidelity simulation for reinforcement learning, yet existing approaches struggle to provide accurate mechanical stress fields within GPU-a…
Learning Human-Intention Priors from Large-Scale Human Demonstrations for Robotic Manipulation
Yifan Xie, YuAn Wang, Guangyu Chen +3
Human videos contain rich manipulation priors, but using them for robot learning remains difficult because raw observations entangle scene understanding, human motion, and embodime…
OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation
Yushan Liu, Peibo Sun, Shoujie Li +7
World Action Models (WAMs) enhance Vision-Language-Action policies by jointly predicting scene evolution and robot actions, but existing methods usually represent the predicted wor…