3 papers
cs.RO2025
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
Yiguo Fan, Pengxiang Ding, Shuanghao Bai +10
Vision-Language-Action (VLA) models have become a cornerstone in robotic policy learning, leveraging large-scale multimodal data for robust and scalable control. However, existing…
cs.RO2025
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
Xinyang Tong, Pengxiang Ding, Yiguo Fan +9
This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) task…
cs.RO2025
Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
Pengxiang Ding, Jianfei Ma, Xinyang Tong +13
This paper addresses the limitations of current humanoid robot control frameworks, which primarily rely on reactive mechanisms and lack autonomous interaction capabilities due to d…