5 papers
From Generated Human Videos to Physically Plausible Robot Trajectories
James Ni, Zekai Wang, Wei Lin +5
Video generation models are rapidly improving in their ability to synthesize human actions in novel contexts, holding the potential to serve as high-level planners for contextual r…
Whole-Body Conditioned Egocentric Video Prediction
Yutong Bai, Danny Tran, Amir Bar +3
We train models to Predict Ego-centric Video from human Actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic po…
Dual-Process Image Generation
Grace Luo, Jonathan Granskog, Aleksander Holynski +1
Prior methods for controlling image generation are limited in their ability to be taught new tasks. In contrast, vision-language models, or VLMs, can learn tasks in-context and pro…
Vector Quantized Feature Fields for Fast 3D Semantic Lifting
George Tang, Aditya Agarwal, Weiqiao Han +2
We generalize lifting to semantic lifting by incorporating per-view masks that indicate relevant pixels for lifting tasks. These masks are determined by querying corresponding mult…
Navigation World Models
Amir Bar, Gaoyue Zhou, Danny Tran +2
Navigation is a fundamental skill of agents with visual-motor capabilities. We introduce a Navigation World Model (NWM), a controllable video generation model that predicts future…