16 papers
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
Haoxiang Ma, Junhao Cai, Xiaoxu Xu +26
Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In pract…
EgoAERO: Learning Dexterous Manipulation from a Single Egocentric Video without Object Assets
Yichen Niu, Haoran Lv, Xinrui Zhang +12
Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometr…
VISTA: Vision-Grounded and Physics-Validated Adaptation of UMI data for VLA Training
Siyuan Yang, Linzheng Guo, Ouyang Lu +10
Universal Manipulation Interface (UMI) enables scalable real-world robot data collection without hardware-specific teleoperation, yet leveraging UMI data to train large-scale Visio…
AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
Yuhua Jiang, Shuang Cheng, Yan Ding +2
Vision-language-action (VLA) models have recently emerged as a powerful paradigm for building generalist robots. However, traditional VLA models that generate actions through flow…
U-ARM : Ultra low-cost general teleoperation interface for robot manipulation
Yanwen Zou, Zhaoye Zhou, Chenyang Shi +4
We propose U-Arm, a low-cost and rapidly adaptable leader-follower teleoperation framework designed to interface with most of commercially available robotic arms. Our system suppor…
Learning 2D Invariant Affordance Knowledge for 3D Affordance Grounding
Xianqiang Gao, Pingrui Zhang, Delin Qu +4
3D Object Affordance Grounding aims to predict the functional regions on a 3D object and has laid the foundation for a wide range of applications in robotics. Recent advances tackl…