most citedChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

2 citations · 2 across the 5 of their papers we have counts for

collaborators

14 papers

cs.RO2025

Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation

Yichen Zhu, Feifei Feng

Robots operating in complex and uncertain environments face considerable challenges. Advanced robotic systems often rely on extensive datasets to learn manipulation tasks. In contr…

cs.RO2025

dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought

Junjie Wen, Minjie Zhu, Jiaming Liu +6

Vision-Language-Action (VLA) models are emerging as a next-generation paradigm for robotics. We introduce dVLA, a diffusion-based VLA that leverages a multimodal chain-of-thought t…

cs.RO2025

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

Zhongyi Zhou, Yichen Zhu, Junjie Wen +2

Vision-language-action (VLA) models have emerged as the next generation of models in robotics. However, despite leveraging powerful pre-trained Vision-Language Models (VLMs), exist…

cs.RO2025

WorldEval: World Model as Real-World Robot Policies Evaluator

Yaxuan Li, Yichen Zhu, Junjie Wen +2

The field of robotics has made significant strides toward developing generalist robot manipulation policies. However, evaluating these policies in real-world scenarios remains time…

cs.RO2025

PointVLA: Injecting the 3D World into Vision-Language-Action Models

Chengmeng Li, Junjie Wen, Yan Peng +3

Vision-Language-Action (VLA) models excel at robotic tasks by leveraging large-scale 2D vision-language pretraining, but their reliance on RGB images limits spatial reasoning criti…

cs.RO2025

ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration

Minjie Zhu, Yichen Zhu, Jinming Li +6

Imitation learning has proven to be highly effective in teaching robots dexterous manipulation skills. However, it typically relies on large amounts of human demonstration data, wh…