collaborators

11 papers

cs.RO2025

Embodiment Transfer Learning for Vision-Language-Action Models

Chengmeng Li, Yaxin Peng

Vision-language-action (VLA) models have significantly advanced robotic learning, enabling training on large-scale, cross-embodiment data and fine-tuning for specific robots. Howev…

cs.RO2025

CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Jinming Li, Yichen Zhu, Zhibin Tang +8

Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving…

cs.RO2025

Diffusion-VLA: Generalizable and Interpretable Robot Foundation Model via Self-Generated Reasoning

Junjie Wen, Minjie Zhu, Yichen Zhu +8

In this paper, we present DiffusionVLA, a novel framework that seamlessly combines the autoregression model with the diffusion model for learning visuomotor policy. Central to our…

cs.CV2025

Efficient Feature Fusion for UAV Object Detection

Xudong Wang, Yaxin Peng, Chaomin Shen

Object detection in unmanned aerial vehicle (UAV) remote sensing images poses significant challenges due to unstable image quality, small object sizes, complex backgrounds, and env…

cs.RO2025

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Junjie Wen, Yichen Zhu, Jinming Li +9

Vision-Language-Action (VLA) models have shown remarkable potential in visuomotor control and instruction comprehension through end-to-end learning processes. However, current VLA…

cs.RO2025

PointVLA: Injecting the 3D World into Vision-Language-Action Models

Chengmeng Li, Junjie Wen, Yan Peng +3

Vision-Language-Action (VLA) models excel at robotic tasks by leveraging large-scale 2D vision-language pretraining, but their reliance on RGB images limits spatial reasoning criti…