Diffusion Models Beat GANs on Image Synthesis
arXiv:2105.05233
Abstract
We show that diffusion models can achieve image sample quality superior to the current state-of-the-art generative models. We achieve this on unconditional image synthesis by finding a better architecture through a series of ablations. For conditional image synthesis, we further improve sample quality with classifier guidance: a simple, compute-efficient method for trading off diversity for fidelity using gradients from a classifier. We achieve an FID of 2.97 on ImageNet 128128, 4.59 on ImageNet 256256, and 7.72 on ImageNet 512512, and we match BigGAN-deep even with as few as 25 forward passes per sample, all while maintaining better coverage of the distribution. Finally, we find that classifier guidance combines well with upsampling diffusion models, further improving FID to 3.94 on ImageNet 256256 and 3.85 on ImageNet 512512. We release our code at https://github.com/openai/guided-diffusion
Added compute requirements, ImageNet 256256 upsampling FID and samples, DDIM guided sampler, fixed typos
References in corpus (15)
- Conditional Generative Adversarial Nets
- Learning Transferable Visual Models From Natural Language Supervision
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Language Models are Few-Shot Learners
- Score-Based Generative Modeling through Stochastic Differential Equations
- Zero-Shot Text-to-Image Generation
- Improved Denoising Diffusion Probabilistic Models
- Jukebox: A Generative Model for Music
- Generating Diverse High-Fidelity Images with VQ-VAE-2
- High-Fidelity Image Generation With Fewer Labels
- RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation
- Knowledge Distillation in Iterative Generative Models for Improved Sampling Speed
- Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images
- Variational Walkback: Learning a Transition Operator as a Stochastic Recurrent Net
- Generating Images with Sparse Representations
Cited by in corpus (13)
- On Fast Sampling of Diffusion Probabilistic Models
- ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis
- Learning to Efficiently Sample from Diffusion Probabilistic Models
- A Variational Perspective on Diffusion-Based Generative Models and Score Matching
- Bilateral Denoising Diffusion Models
- Diffusion Priors In Variational Autoencoders
- Denoising Diffusion Gamma Models
- Moser Flow: Divergence-based Generative Modeling on Manifolds
- The Future is Log-Gaussian: ResNets and Their Infinite-Depth-and-Width Limit at Initialization
- D2C: Diffusion-Denoising Models for Few-shot Conditional Generation
- Instance-Conditioned GAN
- Zero-Shot Translation using Diffusion Models
- Improving Compositionality of Neural Networks by Decoding Representations to Inputs