7 papers
Realtime-VLA FLASH: Speculative Inference Framework for Diffusion-based VLAs
Jiahui Niu, Kefan Gu, Yucheng Zhao +5
Diffusion-based vision-language-action models (dVLAs) are promising for embodied intelligence but are fundamentally limited in real-time deployment by the high latency of full infe…
Dexbotic: Open-Source Vision-Language-Action Toolbox
Bin Xie, Erjin Zhou, Fan Jia +36
In this paper, we present Dexbotic, an open-source Vision-Language-Action (VLA) model toolbox based on PyTorch. It aims to provide a one-stop VLA research service for professionals…
RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies
Adina Yakefu, Bin Xie, Chongyang Xu +34
Testing on real machines is indispensable for robotic control algorithms. In the context of learning-based algorithms, especially VLA models, demand for large-scale evaluation, i.e…
ManiAgent: An Agentic Framework for General Robotic Manipulation
Yi Yang, Kefan Gu, Yuqing Wen +4
While Vision-Language-Action (VLA) models have demonstrated impressive capabilities in robotic manipulation, their performance in complex reasoning and long-horizon task planning i…
IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
Yandu Chen, Kefan Gu, Yuqing Wen +3
Vision-Language-Action (VLA) models leverage pretrained vision-language models (VLMs) to couple perception with robotic control, offering a promising path toward general-purpose em…
LLaDA-VLA: Vision Language Diffusion Action Models
Yuqing Wen, Hebei Li, Kefan Gu +3
The rapid progress of auto-regressive vision-language models (VLMs) has inspired growing interest in vision-language-action models (VLA) for robotic manipulation. Recently, masked…