7 papers · 1 filter
StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation
Xiao Liu, Yuguang Yang, Xi Wang +6
Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the fut…
JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling
Yihan Lin, Jiawei He, Shifeng Bao +6
Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAM…
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
Shuanghao Bai, Jing Lyu, Wanqi Zhou +9
Vision-Language-Action (VLA) models benefit from chain-of-thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representati…
VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models
Zirui Ge, Pengxiang Ding, Baohua Yin +16
Video action models are an appealing foundation for Vision--Language--Action systems because they can learn visual dynamics from large-scale video data and transfer this knowledge…
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
Shuanghao Bai, Dakai Wang, Cheng Chi +8
In robotic manipulation, vision-language-action (VLA) models have emerged as a promising paradigm for learning generalizable and scalable robot policies. Most existing VLA framewor…
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
Shuanghao Bai, Wenxuan Song, Jiayi Chen +15
Embodied intelligence has witnessed remarkable progress in recent years, driven by advances in computer vision, natural language processing, and the rise of large-scale multimodal…