12 papers
AffordanceSAM: Segment Anything Once More in Affordance Grounding
Dengyang Jiang, Zanyi Wang, Hengzhuang Li +7
Building a generalized affordance grounding model to identify actionable regions on objects is vital for real-world applications. Existing methods to train the model can be divided…
PA-BiCoop: A Primary-Auxiliary Cooperative Framework for General Bimanual Manipulation
Bai Qicheng, Wang Ziru, Ma Teli +3
Bimanual manipulation is essential for advanced robotic systems because it offers higher efficiency and flexibility compared to single-arm configurations. However, existing approac…
Perceptive Behavior Foundation Model: Adapting Human Motion Priors to Robot-Centric Terrain
Zifan Wang, Yizhao Li, Teli Ma +5
Humanoid behavior foundation models aim to acquire reusable whole-body control policies from broad human motion priors, enabling a single controller to produce diverse and expressi…
MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation
Jia Zheng, Teli Ma, Yudong Fan +3
World Action Models (WAMs) couple a video dynamics prior to the policy and have shown encouraging results on tabletop manipulation, but iterative denoising over high-dimensional vi…
DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
Teli Ma, Jia Zheng, Zifan Wang +4
Vision-Language-Action (VLA) models have emerged as a promising paradigm for robot learning, but their representations are still largely inherited from static image-text pretrainin…
Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
Jiaming Zhou, Ke Ye, Jiayi Liu +6
The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However…