collaborators

5 papers

cs.RO2025

Conditioning Matters: Training Diffusion Policies is Faster Than You Think

Zibin Dong, Yicheng Liu, Yinchuan Li +2

Diffusion policies have emerged as a mainstream paradigm for building vision-language-action (VLA) models. Although they demonstrate strong robot control capabilities, their traini…

cs.RO2025

EmbodiedMAE: A Unified 3D Multi-Modal Representation for Robot Manipulation

Zibin Dong, Fei Ni, Yifu Yuan +2

We present EmbodiedMAE, a unified 3D multi-modal representation for robot manipulation. Current approaches suffer from significant domain gaps between training datasets and robot m…

cs.RO2025

Few-Shot Vision-Language Action-Incremental Policy Learning

Mingchen Song, Xiang Deng, Guoqiang Zhong +5

Recently, Transformer-based robotic manipulation methods utilize multi-view spatial representations and language instructions to learn robot motion trajectories by leveraging numer…

cs.RO2025

Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulation

Qi Lv, Hao Li, Xiang Deng +6

Despite the significant success of imitation learning in robotic manipulation, its application to bimanual tasks remains highly challenging. Existing approaches mainly learn a poli…

cs.CV2025

3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds

Hengshuo Chu, Xiang Deng, Qi Lv +4

3D Affordance detection is a challenging problem with broad applications on various robotic tasks. Existing methods typically formulate the detection paradigm as a label-based sema…