1 citations · 2 across the 3 of their papers we have counts for
3 papers · 1 filter
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
Yiguo Fan, Pengxiang Ding, Shuanghao Bai +10
Vision-Language-Action (VLA) models have become a cornerstone in robotic policy learning, leveraging large-scale multimodal data for robust and scalable control. However, existing…
Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
Pengxiang Ding, Jianfei Ma, Xinyang Tong +13
This paper addresses the limitations of current humanoid robot control frameworks, which primarily rely on reactive mechanisms and lack autonomous interaction capabilities due to d…
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
Xinyang Tong, Pengxiang Ding, Yiguo Fan +9
This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) task…