3 papers
cs.RO2026
VISTA: Vision-Grounded and Physics-Validated Adaptation of UMI data for VLA Training
Siyuan Yang, Linzheng Guo, Ouyang Lu +10
Universal Manipulation Interface (UMI) enables scalable real-world robot data collection without hardware-specific teleoperation, yet leveraging UMI data to train large-scale Visio…
cs.AI2026
PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations
Yang Zhang, Jiangyuan Zhao, Chenyou Fan +11
Vision-Language-Action (VLA) models advance robotic control via strong visual-linguistic priors. However, existing VLAs predominantly frame pretraining as supervised behavior cloni…
cs.CV2025
Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction
Chenyou Fan, Fangzheng Yan, Chenjia Bai +4
Learning a generalizable bimanual manipulation policy is extremely challenging for embodied agents due to the large action space and the need for coordinated arm movements. Existin…