activity
20242026
collaborators
Showing cs.ROShow all

12 papers · 1 filter

cs.RO2026

4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields

Lishan Yang, Wenxuan Song, Xi Wang +14

Building on recent advances in world models, World Action Models (WAMs) jointly model video prediction and action generation. However, they typically represent videos in 2D pixel s…

cs.RO2026

SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models

Junjie He, Junfeng Li, Zhide Zhong +9

World-Action Models (WAMs) have emerged as a promising paradigm for robotic manipulation. However, most existing WAMs generate future videos and actions by relying mainly on visual…

cs.RO2026

Is Forward Prediction Enough? Physical State Grounding for JEPA World Models

Haodong Yan, Jiaguan Zhu, Mingyuan Jia +12

Learning structured and control-relevant latent representations remains a key challenge for world models. Recent JEPA-based world models learn action-conditioned predictive latent…

cs.RO2026

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

Junfeng Li, Junjie He, Zhide Zhong +12

Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an o…

cs.RO2026

GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

GigaWorld Team, Angen Ye, Angyuan Ma +26

World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physicall…

cs.RO2026

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark

Huashuo Lei, Wenxuan Song, Huarui Zhang +10

Memory is a critical component of robotic intelligence, as robots must rely on past observations and actions to accomplish long-horizon tasks in partially observable environments.…