2 citations · 2 across the 5 of their papers we have counts for
14 papers
Let Me Show You: Learning by Retrieving from Egocentric Video for Robotic Manipulation
Yichen Zhu, Feifei Feng
Robots operating in complex and uncertain environments face considerable challenges. Advanced robotic systems often rely on extensive datasets to learn manipulation tasks. In contr…
dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
Junjie Wen, Minjie Zhu, Jiaming Liu +6
Vision-Language-Action (VLA) models are emerging as a next-generation paradigm for robotics. We introduce dVLA, a diffusion-based VLA that leverages a multimodal chain-of-thought t…
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
Zhongyi Zhou, Yichen Zhu, Junjie Wen +2
Vision-language-action (VLA) models have emerged as the next generation of models in robotics. However, despite leveraging powerful pre-trained Vision-Language Models (VLMs), exist…
WorldEval: World Model as Real-World Robot Policies Evaluator
Yaxuan Li, Yichen Zhu, Junjie Wen +2
The field of robotics has made significant strides toward developing generalist robot manipulation policies. However, evaluating these policies in real-world scenarios remains time…
PointVLA: Injecting the 3D World into Vision-Language-Action Models
Chengmeng Li, Junjie Wen, Yan Peng +3
Vision-Language-Action (VLA) models excel at robotic tasks by leveraging large-scale 2D vision-language pretraining, but their reliance on RGB images limits spatial reasoning criti…
ObjectVLA: End-to-End Open-World Object Manipulation Without Demonstration
Minjie Zhu, Yichen Zhu, Jinming Li +6
Imitation learning has proven to be highly effective in teaching robots dexterous manipulation skills. However, it typically relies on large amounts of human demonstration data, wh…