collaborators

8 papers

cs.RO2026

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Xiaomi Robotics Team, Jun Guo, Piaopiao Jin +31

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulatio…

cs.RO2026

AdaReP:Adaptive Re-Planning under Model Mismatch for Neural World-Model Predictive Control

Yutian Cheng, Xiaojian Ma, Xianhao Wang +6

Neural world models coupled with model predictive control (MPC) replan at every environment step to bound accumulated prediction error, but this incurs substantial computational ov…

cs.CV2026

Seeing Through Fog: Towards Fog-Invariant Action Recognition

Enqi Liu, Liyuan Pan, Zhi Gao +2

Foggy conditions are commonly encountered in real-world applications; however, existing action recognition approaches typically assume favorable weather and high-quality video inpu…

cs.CV2026

DEVIS-GRPO: Unleashing GRPO on Dynamic Extreme View Synthesis

Yi Zuo, Huimin Wu, Lingling Li +3

Trajectory-controlled video generation has become essential for controllable video generation. While current methods perform well under small-view camera motions, they degrade sign…

cs.CV2025

From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes

Tianxu Wang, Zhuofan Zhang, Ziyu Zhu +5

3D visual grounding has made notable progress in localizing objects within complex 3D scenes. However, grounding referring expressions beyond objects in 3D scenes remains unexplore…

cs.RO2025

FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation

Jun Guo, Xiaojian Ma, Yikai Wang +3

This paper investigates training better visual world models for robot manipulation, i.e., models that can predict future visual observations by conditioning on past frames and robo…