6 papers
Addressable Memory for Video World Models
Xindi Wu, Sven Elflein, James Lucas +5
We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames.…
SVAG-Bench: A Large-Scale Benchmark for Multi-Instance Spatio-temporal Video Action Grounding
Tanveer Hannan, Shuaicong Wu, Mark Weber +6
A truly capable AI system must do more than detect objects or recognize activities in isolation. It must form unified, grounded representations of who is acting, what they are doin…
VGG-T: Offline Feed-Forward 3D Reconstruction at Scale
Sven Elflein, Ruilong Li, Sérgio Agostinho +4
We present a scalable 3D reconstruction model that addresses a critical limitation in offline feed-forward methods: their computational and memory requirements grow quadratically w…
Native Segmentation Vision Transformers
Guillem Brasó, Aljoša Ošep, Laura Leal-Taixé
Uniform downsampling remains the de facto standard for reducing spatial resolution in vision backbones. In this work, we propose an alternative design built around a content-aware…
Zero-Shot 4D Lidar Panoptic Segmentation
Yushan Zhang, Aljoša Ošep, Laura Leal-Taixé +1
Zero-shot 4D segmentation and recognition of arbitrary objects in Lidar is crucial for embodied navigation, with applications ranging from streaming perception to semantic mapping…
Lidar Panoptic Segmentation in an Open World
Anirudh S Chakravarthy, Meghana Reddy Ganesina, Peiyun Hu +4
Addressing Lidar Panoptic Segmentation (LPS ) is crucial for safe deployment of autonomous vehicles. LPS aims to recognize and segment lidar points w.r.t. a pre-defined vocabulary…