9 papers
RayTun3R: Online Camera Adaptation in 3D Foundation Models
Daniil Sinitsyn, Nikita Araslanov, Daniel Cremers
Recent 3D foundation models, such as DUSt3R, MASt3R, VGGT, , and Depth Anything 3, provide strong feed-forward depth and pose estimates on pinhole imagery, but degrade sharpl…
Segmenting, Fast and Slow: Real-Time Open-Vocabulary Video Instance Segmentation with Dual-Path Processing
Luca Barsellotti, Martin Sundermeyer, Mattia Segu +5
Object-centric models inspired by DETR have become the dominant paradigm for open-vocabulary video instance segmentation (OV-VIS). While recent efforts have reduced the computation…
Scene-Centric Unsupervised Video Panoptic Segmentation
Christoph Reich, Oliver Hahn, Nikita Araslanov +4
Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task se…
Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners
Nikita Araslanov, Martin Sundermeyer, Hidenobu Matsuki +2
One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack representations that effectiv…
Generative Shape Reconstruction with Geometry-Guided Langevin Dynamics
Linus Härenstam-Nielsen, Dmitrii Pozdeev, Thomas Dagès +2
Reconstructing complete 3D shapes from incomplete or noisy observations is a fundamentally ill-posed problem that requires balancing measurement consistency with shape plausibility…
FlowFeat: Pixel-Dense Embedding of Motion Profiles
Nikita Araslanov, Anna Sonnweber, Daniel Cremers
Dense and versatile image representations underpin the success of virtually all computer vision applications. However, state-of-the-art networks, such as transformers, produce low-…