6 papers · 1 filter
XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?
Yixiang Chen, Jiabing Yang, Yuan Xu +10
Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture p…
BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation
Peiyan Li, Yuze Zhu, Yixiang Chen +10
Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existi…
FlowWAM: Optical Flow as a Unified Action Representation for World Action Models
Yixiang Chen, Peiyan Li, Yuan Xu +13
World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for co…
DIM-WAM: World-Action Modeling with Diverse Historical Event Memory
Kai Wang, Zhaopeng Gu, Yixiang Chen +7
World-action models have shown promising robot-manipulation performance by jointly predicting future visual states and actions. However, existing methods mainly rely on short-term…
Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision
Yuan Xu, Yixiang Chen, Kai Wang +5
Vision-Language-Action (VLA) models have shown strong potential for generalizable robotic manipulation. During fine-tuning, however, action supervision applies equally across all t…
SKIP: Sparse Keyframe Interpolation Paradigm for Efficient Embodied World Models
Ziheng He, Yixiang Chen, Ning Yang +11
Embodied world models have emerged as a promising paradigm in robotics by predicting how robot actions affect the surrounding scene. However, the rollout inference remains computat…