14 papers
TacWAM: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction
Lei Jin, Yiding Ma, Xin Zhang +3
The paper introduces TacWAM, a mechanics-aware tactile world action model that predicts future tactile signals and uses them as supervision for training contact-rich robot manipula…
WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory
Haisheng Su, Zongdai Liu, Xin Jin +13
World Action Models (WAMs) offer a promising paradigm for robotic manipulation by jointly modeling visual state transitions and robot actions. However, existing WAMs are constraine…
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
Shuailei Ma, Jiaqi Liao, Xinyang Wang +24
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inheren…
Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control
Jianjie Fang, Yongyan Xu, Ziyou Wang +13
World models are rapidly becoming a core infrastructure for embodied intelligence and interactive agents: they provide controllable simulators in which agents can perceive, act, fo…
Dreaming when Necessary: Advancing World Action Models with Adaptive Multi-Modal Reasoning
Yinzhou Tang, Jingbo Xu, Yu Shang +4
World Action Models (WAMs) offer a promising approach to embodied intelligence, yet existing methods rely heavily on video prediction as action priors and lack adaptive multimodal…
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
Baining Zhao, Jiacheng Xu, Weicheng Feng +13
Aerial vision-language navigation (VLN) requires agents to follow natural-language instructions through closed-loop perception and action in 3D environments. We argue that aerial V…