4 papers
APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
Kechun Xu, Zhenjie Zhu, Anzhe Chen +2
Vision-Language-Action (VLA) models that couple pretrained Vision-Language Models (VLMs) with continuous action experts have achieved strong manipulation performance, yet generaliz…
Direction Matters: Learning Force Direction Enables Sim-to-Real Contact-Rich Manipulation
Yifei Yang, Anzhe Chen, Zhenjie Zhu +6
Sim-to-real transfer for contact-rich manipulation remains challenging due to the inherent discrepancy in contact dynamics. While existing methods often rely on costly real-world d…
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
Kechun Xu, Zhenjie Zhu, Anzhe Chen +7
The pursuit of out-of-distribution generalization in Vision-Language-Action (VLA) models is often hindered by catastrophic forgetting of the Vision-Language Model (VLM) backbone du…
Toward Embodiment Equivariant Vision-Language-Action Policy
Anzhe Chen, Yifei Yang, Zhenjie Zhu +4
Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel…