From the 1 of 9 linked papers with an AI index.
9 papers
EquiFusion: Kinematics-Agnostic Human Motion Prediction via Equivariant Latent Diffusion
Cecilia Curreli, Florian Hofherr, Dominik Muhle +3
EquiFusion is a latent diffusion model for 3D human motion prediction that does not rely on fixed skeleton kinematics, allowing it to generalize across datasets and handle partial…
VOCA: Visual Odometry with Codec Awareness
Nouri Alexander Hilscher, Mateo de Mayo, Dominik Muhle +2
Camera pose estimation from image streams is a critical component of spatial world models that integrate perception into planning and decision-making. Nearly all Visual Odometry (V…
IPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance Proposals
Markus Gross, Aya Fahmy, Danit Niwattananan +4
Semantic Scene Completion (SSC) has emerged as a pivotal approach for jointly learning scene geometry and semantics, enabling downstream applications such as navigation in mobile r…
Foundations and Models in Modern Computer Vision: Key Building Blocks in Landmark Architectures
Radu-Andrei Bourceanu, Neil De La Fuente, Jan Grimm +5
This report analyzes the evolution of key design patterns in computer vision by examining six influential papers. The analysis begins with foundational architectures for image reco…
Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images
Philipp Wulff, Felix Wimbauer, Dominik Muhle +1
Volumetric scene reconstruction from a single image is crucial for a broad range of applications like autonomous driving and robotics. Recent volumetric reconstruction methods achi…
GECO: Geometrically Consistent Embedding with Lightspeed Inference
Regine Hartwig, Dominik Muhle, Riccardo Marin +1
Recent advances in feature learning have shown that self-supervised vision foundation models can capture semantic correspondences but often lack awareness of underlying 3D geometry…