8 papers
StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation
Xiao Liu, Yuguang Yang, Xi Wang +6
Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the fut…
JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling
Yihan Lin, Jiawei He, Shifeng Bao +6
Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAM…
MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight
Zehua Fan, Junjie He, Wenxuan Song +14
World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands sim…
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
Shuanghao Bai, Jing Lyu, Wanqi Zhou +9
Vision-Language-Action (VLA) models benefit from chain-of-thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representati…
VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models
Zirui Ge, Pengxiang Ding, Baohua Yin +16
Video action models are an appealing foundation for Vision--Language--Action systems because they can learn visual dynamics from large-scale video data and transfer this knowledge…
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
Shuanghao Bai, Dakai Wang, Cheng Chi +8
In robotic manipulation, vision-language-action (VLA) models have emerged as a promising paradigm for learning generalizable and scalable robot policies. Most existing VLA framewor…