collaborators
Showing cs.ROShow all

7 papers · 1 filter

cs.RO2026

StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation

Xiao Liu, Yuguang Yang, Xi Wang +6

Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the fut…

cs.RO2026

JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

Yihan Lin, Jiawei He, Shifeng Bao +6

Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAM…

cs.RO2026

Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models

Shuanghao Bai, Jing Lyu, Wanqi Zhou +9

Vision-Language-Action (VLA) models benefit from chain-of-thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representati…

cs.RO2026

VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models

Zirui Ge, Pengxiang Ding, Baohua Yin +16

Video action models are an appealing foundation for Vision--Language--Action systems because they can learn visual dynamics from large-scale video data and transfer this knowledge…

cs.RO2026

Reshaping Action Error Distributions for Reliable Vision-Language-Action Models

Shuanghao Bai, Dakai Wang, Cheng Chi +8

In robotic manipulation, vision-language-action (VLA) models have emerged as a promising paradigm for learning generalizable and scalable robot policies. Most existing VLA framewor…

cs.RO2025

Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey

Shuanghao Bai, Wenxuan Song, Jiayi Chen +15

Embodied intelligence has witnessed remarkable progress in recent years, driven by advances in computer vision, natural language processing, and the rise of large-scale multimodal…