2 citations · 2 across the 19 of their papers we have counts for
27 papers
GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning
Long Xu, Zhiqi Zhang, Mi Yan +8
Robot manipulation foundation models require scalable evaluation and data generation across diverse scenarios, with simulation providing an environment for both. Automated scene ge…
ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop Manipulation
Mi Yan, Wenhao Zhang, Zhiqi Zhang +14
Zero-shot generalization to unseen embodiments is important for generalizable vision-language-action (VLA) models as robot hardware evolves and task-specific data collection remain…
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
Yusen Feng, Bingchen Han, Jiangran Lyu +13
Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fin…
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
AllDayNav: Lifelong Navigation via Real-World Reinforcement Learning
Hang Yin, Yinan Liang, Jiazhao Zhang +4
Lifelong embodied navigation in dynamic environments requires robots to form persistent scene understanding from fragmentary observations, which remains difficult for existing meth…
AnchorVLA: Anchored Diffusion for Efficient End-to-End Mobile Manipulation
Jia Syuen Lim, Zhizhen Zhang, Peter Bohm +3
A central challenge in mobile manipulation is preserving multiple plausible action models while remaining reactive during execution. A bottle in a cluttered scene can often be appr…