8 papers · 1 filter
Importance-Aware Low-Rank Distillation of Diffusion Transformers
Denis Zavadski, Sebastian Heid, Damjan Kalšan +2
Diffusion Transformers (DiTs) have emerged as a dominant architecture for high-quality text-to-image generation, yet their scale poses challenges for efficient deployment. While tr…
Do Image Editing Models Understand Lighting?
Tim Küchler, Johann-Friedrich Feiden, Matthias Nießner +1
While recent advancements in generative image editing models have achieved stunning visual fidelity, it remains an open question whether these systems possess an intrinsic knowledg…
Online Video Depth Anything: Temporally-Consistent Depth Prediction with Low Memory Consumption
Johann-Friedrich Feiden, Tim Küchler, Denis Zavadski +2
Depth estimation from monocular video has become a key component of many real-world computer vision systems. Recently, Video Depth Anything (VDA) has demonstrated strong performanc…
Product-Quantised Image Representation for High-Quality Image Synthesis
Denis Zavadski, Nikita Philip Tatsch, Carsten Rother
Product quantisation (PQ) is a classical method for scalable vector encoding, yet it has seen limited usage for latent representations in high-fidelity image generation. In this wo…
A Framework for Low-Effort Training Data Generation for Urban Semantic Segmentation
Damjan Kalšan, Denis Zavadski, Tim Küchler +3
Synthetic datasets are widely used for training urban scene recognition models, but even highly realistic renderings show a noticeable gap to real imagery. This gap is particularly…
PrimeDepth: Efficient Monocular Depth Estimation with a Stable Diffusion Preimage
Denis Zavadski, Damjan Kalšan, Carsten Rother
This work addresses the task of zero-shot monocular depth estimation. A recent advance in this field has been the idea of utilising Text-to-Image foundation models, such as Stable…