collaborators

7 papers

cs.CV2026

Addressable Memory for Video World Models

Xindi Wu, Sven Elflein, James Lucas +5

We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames.…

cs.CV2026

Motion Attribution for Video Generation

Xindi Wu, Despoina Paschalidou, Jun Gao +5

Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a m…

cs.CV2026

DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models

Xinglong Sun, Kevin Xie, Jenny Schmalfuss +5

Autonomous driving has shifted towards end-to-end policy learning, where reliable, interpretable policy evaluation is a fundamental challenge as driving quality is highly context-d…

cs.CV2026

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation

NVIDIA, :, Aarti Basant +32

As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. In closed-loop simulation, the driving pol…

cs.CV2025

Cosmos World Foundation Model Platform for Physical AI

NVIDIA, :, Niket Agarwal +76

Physical AI needs to be trained digitally first. It needs a digital twin of itself, the policy model, and a digital twin of the world, the world model. In this paper, we present th…

cs.GR2025

VideoPanda: Video Panoramic Diffusion with Multi-view Attention

Kevin Xie, Amirmojtaba Sabour, Jiahui Huang +5

High resolution panoramic video content is paramount for immersive experiences in Virtual Reality, but is non-trivial to collect as it requires specialized equipment and intricate…