2 citations · 4 across the 21 of their papers we have counts for
13 papers · 1 filter
DualWAM: Dual-System World Action Models for Asynchronous Global Planning and Local Refinement
Yixin Zheng, Jiangran Lyu, Yuntian Deng +6
World Action Models (WAMs) jointly generate robot actions and predict future world states, transferring priors from video pretraining to robot control. However, future visual predi…
GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning
Long Xu, Zhiqi Zhang, Mi Yan +8
Robot manipulation foundation models require scalable evaluation and data generation across diverse scenarios, with simulation providing an environment for both. Automated scene ge…
ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop Manipulation
Mi Yan, Wenhao Zhang, Zhiqi Zhang +14
Zero-shot generalization to unseen embodiments is important for generalizable vision-language-action (VLA) models as robot hardware evolves and task-specific data collection remain…
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
Yusen Feng, Bingchen Han, Jiangran Lyu +13
Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fin…
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
AllDayNav: Lifelong Navigation via Real-World Reinforcement Learning
Hang Yin, Yinan Liang, Jiazhao Zhang +4
Lifelong embodied navigation in dynamic environments requires robots to form persistent scene understanding from fragmentary observations, which remains difficult for existing meth…