5 papers
Addressable Memory for Video World Models
Xindi Wu, Sven Elflein, James Lucas +5
We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames.…
Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation
NVIDIA, :, Jiahui Huang +14
The paper introduces Instant NuRec, a feed‑forward neural model that converts a short multi‑camera driving log into a fully simulatable 3D Gaussian Splatting scene in about 1.5 sec…
Motion Attribution for Video Generation
Xindi Wu, Despoina Paschalidou, Jun Gao +5
Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a m…
VGG-T: Offline Feed-Forward 3D Reconstruction at Scale
Sven Elflein, Ruilong Li, Sérgio Agostinho +4
We present a scalable 3D reconstruction model that addresses a critical limitation in offline feed-forward methods: their computational and memory requirements grow quadratically w…
Depth Completion as Parameter-Efficient Test-Time Adaptation
Bingxin Ke, Qunjie Zhou, Jiahui Huang +5
We introduce CAPA, a parameter-efficient test-time optimization framework that adapts pre-trained 3D foundation models (FMs) for depth completion, using sparse geometric cues. Unli…