8 papers
Déjà View: Looping Transformers for Multi-View 3D Reconstruction
Alessandro Burzio, Tobias Fischer, Sven Elflein +9
Recent feed-forward 3D reconstruction transformers have scaled to over a billion parameters, following the broader trend of increasing model capacity in computer vision. Yet emergi…
VGG-T: Offline Feed-Forward 3D Reconstruction at Scale
Sven Elflein, Ruilong Li, Sérgio Agostinho +4
We present a scalable 3D reconstruction model that addresses a critical limitation in offline feed-forward methods: their computational and memory requirements grow quadratically w…
Depth Completion as Parameter-Efficient Test-Time Adaptation
Bingxin Ke, Qunjie Zhou, Jiahui Huang +5
We introduce CAPA, a parameter-efficient test-time optimization framework that adapts pre-trained 3D foundation models (FMs) for depth completion, using sparse geometric cues. Unli…
A Guide to Structureless Visual Localization
Vojtech Panek, Qunjie Zhou, Yaqing Ding +4
Visual localization algorithms, i.e., methods that estimate the camera pose of a query image in a known scene, are core components of many applications, including self-driving cars…
DynOMo: Online Point Tracking by Dynamic Online Monocular Gaussian Reconstruction
Jenny Seidenschwarz, Qunjie Zhou, Bardienus Duisterhof +2
Reconstructing scenes and tracking motion are two sides of the same coin. Tracking points allow for geometric reconstruction [14], while geometric reconstruction of (dynamic) scene…
MATCHA:Towards Matching Anything
Fei Xue, Sven Elflein, Laura Leal-Taixé +1
Establishing correspondences across images is a fundamental challenge in computer vision, underpinning tasks like Structure-from-Motion, image editing, and point tracking. Traditio…