9 papers
MMDiff: Extending Diffusion Transformers for Multi-Modal Generation
Yagmur Akarken, Orest Kupyn, Christian Rupprecht
Diffusion transformers have demonstrated remarkable generative capabilities, yet the rich perceptual representations computed across their denoising trajectory are discarded once t…
FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation
Orest Kupyn, Goutam Bhat, Philipp Henzler +3
Generating explorable 3D scenes from a single image requires strong generative priors and accurate geometric representations suitable for downstream use. Current video diffusion mo…
Epipolar Geometry Improves Video Generation Models
Orest Kupyn, Théo Uscidda, Marta Tintore Gazulla +3
Video generation models have advanced significantly through the latent diffusion transformers trained with rectified flow techniques. Yet these models still struggle with geometric…
S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
Orest Kupyn, Hirokatsu Kataoka, Christian Rupprecht
Salient object detection exemplifies data-bounded tasks where expensive pixel-precise annotations force separate model training for related subtasks like DIS and HR-SOD. We present…
VGGHeads: 3D Multi Head Alignment with a Large-Scale Synthetic Dataset
Orest Kupyn, Eugene Khvedchenia, Christian Rupprecht
Human head detection, keypoint estimation, and 3D head model fitting are essential tasks with many applications. However, traditional real-world datasets often suffer from bias, pr…
Stepper: Stepwise Immersive Scene Generation with Multiview Panoramas
Felix Wimbauer, Fabian Manhardt, Michael Oechsle +4
The synthesis of immersive 3D scenes from text is rapidly maturing, driven by novel video generative models and feed-forward 3D reconstruction, with vast potential in AR/VR and wor…