6 papers
Rethinking Pixel Mean Flows via Interval Denoiser
Alexander Zaytsev, Dmitry Baranchuk, Alexander Korotin +1
Modern diffusion and flow-based models are increasingly moving toward few-step, latent-free generation to bypass the computational overhead of multi-step sampling and the reconstru…
Alchemist: Turning Public Text-to-Image Data into Generative Gold
Valerii Startsev, Alexander Ustyuzhanin, Alexey Kirillov +2
Pre-training equips text-to-image (T2I) models with broad world knowledge, but this alone is often insufficient to achieve high aesthetic quality and alignment. Consequently, super…
CasTex: Cascaded Text-to-Texture Synthesis via Explicit Texture Maps and Physically-Based Shading
Mishan Aliev, Dmitry Baranchuk, Kirill Struminsky
This work investigates text-to-texture synthesis using diffusion models to generate physically-based texture maps. We aim to achieve realistic model appearances under varying light…
MADrive: Memory-Augmented Driving Scene Modeling
Polina Karpikova, Daniil Selikhanovych, Kirill Struminsky +3
Recent advances in scene reconstruction have pushed toward highly realistic modeling of autonomous driving (AD) environments using 3D Gaussian splatting. However, the resulting rec…
Inverse Bridge Matching Distillation
Nikita Gushchin, David Li, Daniil Selikhanovych +3
Learning diffusion bridge models is easy; making them fast and practical is an art. Diffusion bridge models (DBMs) are a promising extension of diffusion models for applications in…
Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis
Anton Voronov, Denis Kuznedelev, Mikhail Khoroshikh +2
This work presents Switti, a scale-wise transformer for text-to-image generation. We start by adapting an existing next-scale prediction autoregressive (AR) architecture to T2I gen…