10 papers
UniMM: A Unified Mixture Model Framework for Multi-Agent Simulation
Longzhong Lin, Xuewu Lin, Kechun Xu +4
Simulation plays a crucial role in assessing autonomous driving systems, where the generation of realistic multi-agent behaviors is a key aspect. In multi-agent simulation, the pri…
APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
Kechun Xu, Zhenjie Zhu, Anzhe Chen +2
Vision-Language-Action (VLA) models that couple pretrained Vision-Language Models (VLMs) with continuous action experts have achieved strong manipulation performance, yet generaliz…
Learning A Simulation-based Visual Policy for Real-world Peg In Unseen Holes
Liang Xie, Hongxiang Yu, Kechun Xu +5
This paper proposes a learning-based visual peg-in-hole that enables training with several shapes in simulation, and adapting to arbitrary unseen shapes in real world with minimal…
Direction Matters: Learning Force Direction Enables Sim-to-Real Contact-Rich Manipulation
Yifei Yang, Anzhe Chen, Zhenjie Zhu +6
Sim-to-real transfer for contact-rich manipulation remains challenging due to the inherent discrepancy in contact dynamics. While existing methods often rely on costly real-world d…
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
Kechun Xu, Zhenjie Zhu, Anzhe Chen +7
The pursuit of out-of-distribution generalization in Vision-Language-Action (VLA) models is often hindered by catastrophic forgetting of the Vision-Language Model (VLM) backbone du…
Toward Embodiment Equivariant Vision-Language-Action Policy
Anzhe Chen, Yifei Yang, Zhenjie Zhu +4
Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel…