5 papers · 1 filter
SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models
Junjie He, Junfeng Li, Zhide Zhong +9
World-Action Models (WAMs) have emerged as a promising paradigm for robotic manipulation. However, most existing WAMs generate future videos and actions by relying mainly on visual…
Is Forward Prediction Enough? Physical State Grounding for JEPA World Models
Haodong Yan, Jiaguan Zhu, Mingyuan Jia +12
Learning structured and control-relevant latent representations remains a key challenge for world models. Recent JEPA-based world models learn action-conditioned predictive latent…
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation
Junfeng Li, Junjie He, Zhide Zhong +12
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an o…
VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation
Zhide Zhong, Haodong Yan, Junfeng Li +3
Although pre-trained Vision-Language-Action (VLA) models exhibit impressive generalization in robotic manipulation, post-training remains crucial to ensure reliable performance dur…
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
Zhide Zhong, Haodong Yan, Junfeng Li +8
Many Vision-Language-Action (VLA) models are built upon an internal world model trained via next-frame prediction ``''. However, this paradigm attempts to…