collaborators

6 papers

cs.RO2026

APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies

Kechun Xu, Zhenjie Zhu, Anzhe Chen +2

Vision-Language-Action (VLA) models that couple pretrained Vision-Language Models (VLMs) with continuous action experts have achieved strong manipulation performance, yet generaliz…

cs.RO2026

D-REX: Differentiable Real-to-Sim-to-Real Engine for Learning Dexterous Grasping

Haozhe Lou, Mingtong Zhang, Haoran Geng +9

Simulation provides a cost-effective and flexible platform for data generation and policy learning to develop robotic systems. However, bridging the gap between simulation and real…

cs.RO2025

Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy

Kechun Xu, Zhenjie Zhu, Anzhe Chen +7

The pursuit of out-of-distribution generalization in Vision-Language-Action (VLA) models is often hindered by catastrophic forgetting of the Vision-Language Model (VLM) backbone du…

cs.RO2025

Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter

Kechun Xu, Xunlong Xia, Kaixuan Wang +6

We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches le…

cs.CV2025

Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions

He Zhu, Quyu Kong, Kechun Xu +5

Grounding 3D object affordance is a task that locates objects in 3D space where they can be manipulated, which links perception and action for embodied intelligence. For example, f…

cs.RO2025

Grasp, See, and Place: Efficient Unknown Object Rearrangement with Policy Structure Prior

Kechun Xu, Zhongxiang Zhou, Jun Wu +3

We focus on the task of unknown object rearrangement, where a robot is supposed to re-configure the objects into a desired goal configuration specified by an RGB-D image. Recent wo…