Elucidating the Design Space of Diffusion-Based Generative Models
arXiv:2206.00364
Abstract
We argue that the theory and practice of diffusion-based generative models are currently unnecessarily convoluted and seek to remedy the situation by presenting a design space that clearly separates the concrete design choices. This lets us identify several changes to both the sampling and training processes, as well as preconditioning of the score networks. Together, our improvements yield new state-of-the-art FID of 1.79 for CIFAR-10 in a class-conditional setting and 1.97 in an unconditional setting, with much faster sampling (35 network evaluations per image) than prior designs. To further demonstrate their modular nature, we show that our design changes dramatically improve both the efficiency and quality obtainable with pre-trained score networks from previous work, including improving the FID of a previously trained ImageNet-64 model from 2.07 to near-SOTA 1.55, and after re-training with our proposed improvements to a new SOTA of 1.36.
NeurIPS 2022
Cited by in corpus (26)
- Diffusion Models in Vision: A Survey
- Speech Enhancement and Dereverberation with Diffusion-based Generative Models
- DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models
- Diffusion Models, Image Super-Resolution And Everything: A Survey
- LDMVFI: Video Frame Interpolation with Latent Diffusion Models
- Generative Discovery of Novel Chemical Designs using Diffusion Modeling and Transformer Deep Neural Networks with Application to Deep Eutectic Solvents
- PC-JeDi: Diffusion for Particle Cloud Generation in High Energy Physics
- Deep learning probability flows and entropy production rates in active matter
- Deep Generative Models for Detector Signature Simulation: A Taxonomic Review
- StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
- Face Morphing Attack Detection with Denoising Diffusion Probabilistic Models
- Iterative -(de)Blending: a Minimalist Deterministic Diffusion Model
- A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
- Diffusion-Based Audio Inpainting
- Surgical Text-to-Image Generation
- SEEDS: Exponential SDE Solvers for Fast High-Quality Sampling from Diffusion Models
- RGB-D-Fusion: Image Conditioned Depth Diffusion of Humanoid Subjects
- Optimal Linear Subspace Search: Learning to Construct Fast and High-Quality Schedulers for Diffusion Models
- Back-Projection Diffusion: Solving the Wideband Inverse Scattering Problem with Diffusion Models
- Fast Inference in Denoising Diffusion Models via MMD Finetuning
- Zero Shot Molecular Generation via Similarity Kernels
- Attacking the Spike: On the Transferability and Security of Spiking Neural Networks to Adversarial Examples
- CoSTI: Consistency Models for (a faster) Spatio-Temporal Imputation
- Hand-Shadow Poser
- Turbulent Injection assisted by Diffusion Models for Scale Resolving Simulations
- U-Turn Diffusion