7 papers
Addressable Memory for Video World Models
Xindi Wu, Sven Elflein, James Lucas +5
We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames.…
FlowBender: Feedback-Aware Training for Self-Correcting Conditional Flows
Daniel Gilo, Sven Elflein, Ido Sobol +1
Conditional diffusion and flow models routinely fail to satisfy the very constraints that define their task. For instance, a depth-conditioned model often produces images whose re-…
Déjà View: Looping Transformers for Multi-View 3D Reconstruction
Alessandro Burzio, Tobias Fischer, Sven Elflein +9
Recent feed-forward 3D reconstruction transformers have scaled to over a billion parameters, following the broader trend of increasing model capacity in computer vision. Yet emergi…
VGG-T: Offline Feed-Forward 3D Reconstruction at Scale
Sven Elflein, Ruilong Li, Sérgio Agostinho +4
We present a scalable 3D reconstruction model that addresses a critical limitation in offline feed-forward methods: their computational and memory requirements grow quadratically w…
LongPerceptualThoughts: Distilling System-2 Reasoning for System-1 Perception
Yuan-Hong Liao, Sven Elflein, Liu He +4
Recent reasoning models through test-time scaling have demonstrated that long chain-of-thoughts can unlock substantial performance boosts in hard reasoning tasks such as math and c…
MATCHA:Towards Matching Anything
Fei Xue, Sven Elflein, Laura Leal-Taixé +1
Establishing correspondences across images is a fundamental challenge in computer vision, underpinning tasks like Structure-from-Motion, image editing, and point tracking. Traditio…