collaborators

5 papers

cs.RO2026

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models

Junjie He, Junfeng Li, Zhide Zhong +9

World-Action Models (WAMs) have emerged as a promising paradigm for robotic manipulation. However, most existing WAMs generate future videos and actions by relying mainly on visual…

cs.RO2026

Is Forward Prediction Enough? Physical State Grounding for JEPA World Models

Haodong Yan, Jiaguan Zhu, Mingyuan Jia +12

Learning structured and control-relevant latent representations remains a key challenge for world models. Recent JEPA-based world models learn action-conditioned predictive latent…

cs.CV2026

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

Haodong Yan, Junfeng Li, Junjie He +12

Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, transferring their learned dynamics prior for action prediction. These VGMs…

cs.RO2026

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

Junfeng Li, Junjie He, Zhide Zhong +12

Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an o…

cs.CV2026

S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight

Haodong Yan, Zhide Zhong, Jiaguan Zhu +10

Video action models (VAMs) have emerged as a promising paradigm for robot learning, owing to their powerful visual foresight for complex manipulation tasks. However, current VAMs,…