collaborators

19 papers

cs.CV2026

PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

Haofei Xu, Rundi Wu, Philipp Henzler +7

State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions, or compress geometry into latent spaces in order to leverage…

cs.CV2026

Unified Semantic Transformer for 3D Scene Understanding

Sebastian Koch, Johanna Wald, Hidenobu Matsuki +3

Holistic 3D scene understanding involves capturing and parsing unstructured 3D environments. Due to the inherent complexity of the real world, existing models have predominantly be…

cs.CV2026

OpenGaFF: Open-Vocabulary Gaussian Feature Field with Codebook Attention

Kunyi Li, Michael Niemeyer, Sen Wang +3

Understanding open-vocabulary 3D scenes with Gaussian-based representations remains challenging due to fragmented and spatially inconsistent semantic predictions across multi-view…

cs.CV2026

Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas

Felix Wimbauer, Fabian Manhardt, Michael Oechsle +4

The synthesis of immersive 3D scenes from text is rapidly maturing, driven by novel video generative models and feed-forward 3D reconstruction, with vast potential in AR/VR and wor…

cs.GR2025

3D-LATTE: Latent Space 3D Editing from Textual Instructions

Maria Parelli, Michael Oechsle, Michael Niemeyer +2

Despite the recent success of multi-view diffusion models for text/image-based 3D asset generation, instruction-based editing of 3D assets lacks surprisingly far behind the quality…

cs.CV2025

SegSplat: Feed-forward Gaussian Splatting and Open-Set Semantic Segmentation

Peter Siegel, Federico Tombari, Marc Pollefeys +1

We have introduced SegSplat, a novel framework designed to bridge the gap between rapid, feed-forward 3D reconstruction and rich, open-vocabulary semantic understanding. By constru…