collaborators

8 papers

cs.CV2026

Lifting Embodied World Models for Planning and Control

Alex N. Wang, Trevor Darrell, Pavel Izmailov +2

World models of embodied agents predict future observations conditioned on an action taken by the agent. For complex embodiments, action spaces are high-dimensional and difficult t…

cs.RO2025

From Generated Human Videos to Physically Plausible Robot Trajectories

James Ni, Zekai Wang, Wei Lin +5

Video generation models are rapidly improving in their ability to synthesize human actions in novel contexts, holding the potential to serve as high-level planners for contextual r…

cs.CV2025

Whole-Body Conditioned Egocentric Video Prediction

Yutong Bai, Danny Tran, Amir Bar +3

We train models to Predict Ego-centric Video from human Actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic po…

cs.CV2025

Dual-Process Image Generation

Grace Luo, Jonathan Granskog, Aleksander Holynski +1

Prior methods for controlling image generation are limited in their ability to be taught new tasks. In contrast, vision-language models, or VLMs, can learn tasks in-context and pro…

cs.CV2025

Vision-Language Models Create Cross-Modal Task Representations

Grace Luo, Trevor Darrell, Amir Bar

Autoregressive vision-language models (VLMs) can handle many tasks within a single model, yet the representations that enable this capability remain opaque. We find that VLMs align…

cs.CV2025

Navigation World Models

Amir Bar, Gaoyue Zhou, Danny Tran +2

Navigation is a fundamental skill of agents with visual-motor capabilities. We introduce a Navigation World Model (NWM), a controllable video generation model that predicts future…