11 papers
Embodiment Transfer Learning for Vision-Language-Action Models
Chengmeng Li, Yaxin Peng
Vision-language-action (VLA) models have significantly advanced robotic learning, enabling training on large-scale, cross-embodiment data and fine-tuning for specific robots. Howev…
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
Jinming Li, Yichen Zhu, Zhibin Tang +8
Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving…
Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning
Junjie Wen, Minjie Zhu, Yichen Zhu +8
In this paper, we present DiffusionVLA, a novel framework that seamlessly combines the autoregression model with the diffusion model for learning visuomotor policy. Central to our…
Efficient Feature Fusion for UAV Object Detection
Xudong Wang, Yaxin Peng, Chaomin Shen
Object detection in unmanned aerial vehicle (UAV) remote sensing images poses significant challenges due to unstable image quality, small object sizes, complex backgrounds, and env…
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
Junjie Wen, Yichen Zhu, Jinming Li +9
Vision-Language-Action (VLA) models have shown remarkable potential in visuomotor control and instruction comprehension through end-to-end learning processes. However, current VLA…
PointVLA: Injecting the 3D World into Vision-Language-Action Models
Chengmeng Li, Junjie Wen, Yan Peng +3
Vision-Language-Action (VLA) models excel at robotic tasks by leveraging large-scale 2D vision-language pretraining, but their reliance on RGB images limits spatial reasoning criti…