From the 1 of 11 linked papers with an AI index.
11 papers
RoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation
Byeongguk Jeon, Seonghyeon Ye, JaeHyeok Doo +4
RoboWorld is an automated pipeline that uses a fast autoregressive video world model and a vision-language scoring system to evaluate generalist robot policies efficiently and reli…
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy
JaeHyeok Doo, Byeongguk Jeon, Seonghyeon Ye +2
There is growing interest in utilizing flow-based models as decision-making policies in reinforcement learning due to their high expressive capacity. However, effectively leveragin…
World Simulation with Video Foundation Models for Physical AI
NVIDIA, :, Arslan Ali +87
We introduce [Cosmos-Predict2.5], the latest generation of the Cosmos World Foundation Models for Physical AI. Built on a flow-based architecture, [Cosmos-Predict2.5] unifies Text2…
World Action Models are Zero-shot Policies
Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng +33
State-of-the-art Vision-Language-Action (VLA) models excel at semantic generalization but struggle to generalize to unseen physical motions in novel environments. We introduce Drea…
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
Shenyuan Gao, William Liang, Kaiyuan Zheng +27
Being able to simulate the outcomes of actions in varied environments will revolutionize the development of generalist agents at scale. However, modeling these world dynamics, espe…
DreamGen: Unlocking Generalization in Robot Learning through Video World Models
Joel Jang, Seonghyeon Ye, Zongyu Lin +25
We introduce DreamGen, a simple yet highly effective 4-stage pipeline for training robot policies that generalize across behaviors and environments through neural trajectories - sy…