8 papers
RayTun3R: Online Camera Adaptation in 3D Foundation Models
Daniil Sinitsyn, Nikita Araslanov, Daniel Cremers
Recent 3D foundation models, such as DUSt3R, MASt3R, VGGT, , and Depth Anything 3, provide strong feed-forward depth and pose estimates on pinhole imagery, but degrade sharpl…
Scene-Centric Unsupervised Video Panoptic Segmentation
Christoph Reich, Oliver Hahn, Nikita Araslanov +4
Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task se…
INSID3: Training-Free In-Context Segmentation with DINOv3
Claudia Cuttano, Gabriele Trivigno, Christoph Reich +3
In-context segmentation (ICS) aims to segment arbitrary concepts, e.g., objects, parts, or personalized instances, given one annotated visual examples. Existing work relies on (i)…
RINO: Rotation-Invariant Non-Rigid Correspondences
Maolin Gao, Shao Jie Hu-Chen, Congyue Deng +3
Dense 3D shape correspondence remains a central challenge in computer vision and graphics as many deep learning approaches still rely on intermediate geometric features or handcraf…
ControlEvents: Controllable Synthesis of Event Camera Datawith Foundational Prior from Image Diffusion Models
Yixuan Hu, Yuxuan Xue, Simon Klenk +2
In recent years, event cameras have gained significant attention due to their bio-inspired properties, such as high temporal resolution and high dynamic range. However, obtaining l…
Motion-Refined DINOSAUR for Unsupervised Multi-Object Discovery
Xinrui Gong, Oliver Hahn, Christoph Reich +4
Unsupervised multi-object discovery (MOD) aims to detect and localize distinct object instances in visual scenes without any form of human supervision. Recent approaches leverage o…