5 papers
Embodiment Transfer Learning for Vision-Language-Action Models
Chengmeng Li, Yaxin Peng
Vision-language-action (VLA) models have significantly advanced robotic learning, enabling training on large-scale, cross-embodiment data and fine-tuning for specific robots. Howev…
ActiveUMI: Robotic Manipulation with Active Perception from Robot-Free Human Demonstrations
Qiyuan Zeng, Chengmeng Li, Jude St. John +5
We present ActiveUMI, a framework for a data collection system that transfers in-the-wild human demonstrations to robots capable of complex bimanual manipulation. ActiveUMI couples…
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
Jinming Li, Yichen Zhu, Zhibin Tang +8
Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving…
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
Junjie Wen, Minjie Zhu, Yichen Zhu +8
In this paper, we present DiffusionVLA, a novel framework that seamlessly combines the autoregression model with the diffusion model for learning visuomotor policy. Central to our…
PointVLA: Injecting the 3D World into Vision-Language-Action Models
Chengmeng Li, Junjie Wen, Yan Peng +3
Vision-Language-Action (VLA) models excel at robotic tasks by leveraging large-scale 2D vision-language pretraining, but their reliance on RGB images limits spatial reasoning criti…