7 citations · 9 across the 12 of their papers we have counts for
16 papers
Robotic Scene Cloning:Advancing Zero-Shot Robotic Scene Adaptation in Manipulation via Visual Prompt Editing
Binyuan Huang, Yuqing Wen, Yucheng Zhao +5
Modern robots can perform a wide range of simple tasks and adapt to diverse scenarios in the well-trained environment. However, deploying pre-trained robot models in real-world use…
ManiAgent: An Agentic Framework for General Robotic Manipulation
Yi Yang, Kefan Gu, Yuqing Wen +4
While Vision-Language-Action (VLA) models have demonstrated impressive capabilities in robotic manipulation, their performance in complex reasoning and long-horizon task planning i…
IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
Yandu Chen, Kefan Gu, Yuqing Wen +3
Vision-Language-Action (VLA) models leverage pretrained vision-language models (VLMs) to couple perception with robotic control, offering a promising path toward general-purpose em…
LLaDA-VLA: Vision Language Diffusion Action Models
Yuqing Wen, Hebei Li, Kefan Gu +3
The rapid progress of auto-regressive vision-language models (VLMs) has inspired growing interest in vision-language-action models (VLA) for robotic manipulation. Recently, masked…
Hita: Holistic Tokenizer for Autoregressive Image Generation
Anlin Zheng, Haochen Wang, Yucheng Zhao +4
Vanilla autoregressive image generation models generate visual tokens step-by-step, limiting their ability to capture holistic relationships among token sequences. Moreover, becaus…
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
Yuqing Wen, Kefan Gu, Haoxuan Liu +4
Vision-Language-Action (VLA) models have recently made significant advance in multi-task, end-to-end robotic control, due to the strong generalization capabilities of Vision-Langua…