activity
20232026
most citedADriver-I: A General World Model for Autonomous Driving

7 citations · 9 across the 12 of their papers we have counts for

collaborators

16 papers

cs.RO2026

Robotic Scene Cloning:Advancing Zero-Shot Robotic Scene Adaptation in Manipulation via Visual Prompt Editing

Binyuan Huang, Yuqing Wen, Yucheng Zhao +5

Modern robots can perform a wide range of simple tasks and adapt to diverse scenarios in the well-trained environment. However, deploying pre-trained robot models in real-world use…

cs.RO2025

ManiAgent: An Agentic Framework for General Robotic Manipulation

Yi Yang, Kefan Gu, Yuqing Wen +4

While Vision-Language-Action (VLA) models have demonstrated impressive capabilities in robotic manipulation, their performance in complex reasoning and long-horizon task planning i…

cs.RO2025

IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction

Yandu Chen, Kefan Gu, Yuqing Wen +3

Vision-Language-Action (VLA) models leverage pretrained vision-language models (VLMs) to couple perception with robotic control, offering a promising path toward general-purpose em…

cs.RO2025

LLaDA-VLA: Vision Language Diffusion Action Models

Yuqing Wen, Hebei Li, Kefan Gu +3

The rapid progress of auto-regressive vision-language models (VLMs) has inspired growing interest in vision-language-action models (VLA) for robotic manipulation. Recently, masked…

cs.CV2025

Hita: Holistic Tokenizer for Autoregressive Image Generation

Anlin Zheng, Haochen Wang, Yucheng Zhao +4

Vanilla autoregressive image generation models generate visual tokens step-by-step, limiting their ability to capture holistic relationships among token sequences. Moreover, becaus…

cs.RO2025

ROSA: Harnessing Robot States for Vision-Language and Action Alignment

Yuqing Wen, Kefan Gu, Haoxuan Liu +4

Vision-Language-Action (VLA) models have recently made significant advance in multi-task, end-to-end robotic control, due to the strong generalization capabilities of Vision-Langua…