8 papers
KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation
Yaping Li, Zhaxizhuoma, Qiaojun Yu +3
Articulated object manipulation requires an understanding of kinematic structure that is difficult and costly to learn from robot demonstrations alone. We introduce the Kinematic-A…
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
Haoxiang Ma, Junhao Cai, Xiaoxu Xu +26
Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In pract…
UMI-Bench 1.0: An Open and Reproducible Real-World Benchmark for Tabletop Robotic Manipulation with UMI Data
Shi Jin, Yuntian Wang, Yuhui Duan +16
Real-robot evaluation is essential for understanding whether learned manipulation policies can operate reliably outside curated demonstrations. This need is particularly pressing f…
EgoAERO: Learning Dexterous Manipulation from a Single Egocentric Video without Object Assets
Yichen Niu, Haoran Lv, Xinrui Zhang +12
Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometr…
VISTA: Vision-Grounded and Physics-Validated Adaptation of UMI data for VLA Training
Siyuan Yang, Linzheng Guo, Ouyang Lu +10
Universal Manipulation Interface (UMI) enables scalable real-world robot data collection without hardware-specific teleoperation, yet leveraging UMI data to train large-scale Visio…
ForceVLA2: Unleashing Hybrid Force-Position Control with Force Awareness for Contact-Rich Manipulation
Yang Li, Zhaxizhuoma, Hongru Jiang +11
Embodied intelligence for contact-rich manipulation has predominantly relied on position control, while explicit awareness and regulation of interaction forces remain under-explore…