7 papers · 1 filter
Lifting Embodied World Models for Planning and Control
Alex N. Wang, Trevor Darrell, Pavel Izmailov +2
World models of embodied agents predict future observations conditioned on an action taken by the agent. For complex embodiments, action spaces are high-dimensional and difficult t…
Whole-Body Conditioned Egocentric Video Prediction
Yutong Bai, Danny Tran, Amir Bar +3
We train models to Predict Ego-centric Video from human Actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic po…
Dual-Process Image Generation
Grace Luo, Jonathan Granskog, Aleksander Holynski +1
Prior methods for controlling image generation are limited in their ability to be taught new tasks. In contrast, vision-language models, or VLMs, can learn tasks in-context and pro…
Vision-Language Models Create Cross-Modal Task Representations
Grace Luo, Trevor Darrell, Amir Bar
Autoregressive vision-language models (VLMs) can handle many tasks within a single model, yet the representations that enable this capability remain opaque. We find that VLMs align…
Navigation World Models
Amir Bar, Gaoyue Zhou, Danny Tran +2
Navigation is a fundamental skill of agents with visual-motor capabilities. We introduce a Navigation World Model (NWM), a controllable video generation model that predicts future…
Vector Quantized Feature Fields for Fast 3D Semantic Lifting
George Tang, Aditya Agarwal, Weiqiao Han +2
We generalize lifting to semantic lifting by incorporating per-view masks that indicate relevant pixels for lifting tasks. These masks are determined by querying corresponding mult…