19 papers
PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation
Haofei Xu, Rundi Wu, Philipp Henzler +7
State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions, or compress geometry into latent spaces in order to leverage…
Unified Semantic Transformer for 3D Scene Understanding
Sebastian Koch, Johanna Wald, Hidenobu Matsuki +3
Holistic 3D scene understanding involves capturing and parsing unstructured 3D environments. Due to the inherent complexity of the real world, existing models have predominantly be…
OpenGaFF: Open-Vocabulary Gaussian Feature Field with Codebook Attention
Kunyi Li, Michael Niemeyer, Sen Wang +3
Understanding open-vocabulary 3D scenes with Gaussian-based representations remains challenging due to fragmented and spatially inconsistent semantic predictions across multi-view…
Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas
Felix Wimbauer, Fabian Manhardt, Michael Oechsle +4
The synthesis of immersive 3D scenes from text is rapidly maturing, driven by novel video generative models and feed-forward 3D reconstruction, with vast potential in AR/VR and wor…
3D-LATTE: Latent Space 3D Editing from Textual Instructions
Maria Parelli, Michael Oechsle, Michael Niemeyer +2
Despite the recent success of multi-view diffusion models for text/image-based 3D asset generation, instruction-based editing of 3D assets lacks surprisingly far behind the quality…
SegSplat: Feed-forward Gaussian Splatting and Open-Set Semantic Segmentation
Peter Siegel, Federico Tombari, Marc Pollefeys +1
We have introduced SegSplat, a novel framework designed to bridge the gap between rapid, feed-forward 3D reconstruction and rich, open-vocabulary semantic understanding. By constru…