Cascaded Diffusion Models for High Fidelity Image Generation
arXiv:2106.15282
Abstract
We show that cascaded diffusion models are capable of generating high fidelity images on the class-conditional ImageNet generation benchmark, without any assistance from auxiliary image classifiers to boost sample quality. A cascaded diffusion model comprises a pipeline of multiple diffusion models that generate images of increasing resolution, beginning with a standard diffusion model at the lowest resolution, followed by one or more super-resolution diffusion models that successively upsample the image and add higher resolution details. We find that the sample quality of a cascading pipeline relies crucially on conditioning augmentation, our proposed method of data augmentation of the lower resolution conditioning inputs to the super-resolution models. Our experiments show that conditioning augmentation prevents compounding error during sampling in a cascaded model, helping us to train cascading pipelines achieving FID scores of 1.48 at 64x64, 3.52 at 128x128 and 4.88 at 256x256 resolutions, outperforming BigGAN-deep, and classification accuracy scores of 63.02% (top-1) and 84.06% (top-5) at 256x256, outperforming VQ-VAE-2.
References in corpus (15)
- Auto-Encoding Variational Bayes
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Large Scale GAN Training for High Fidelity Natural Image Synthesis
- Neural Discrete Representation Learning
- Diffusion Models Beat GANs on Image Synthesis
- Conditional Image Generation with PixelCNN Decoders
- Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
- Pixel Recurrent Neural Networks
- Generative Modeling by Estimating Gradients of the Data Distribution
- Flow++: Improving Flow-Based Generative Models with Variational Dequantization and Architecture Design
- Classification Accuracy Score for Conditional Generative Models
- Improved Techniques for Training Score-Based Generative Models
- Generating High Fidelity Images with Subscale Pixel Networks and Multidimensional Upscaling
- LOGAN: Latent Optimisation for Generative Adversarial Networks
- Hierarchical Autoregressive Image Models with Auxiliary Decoders
Cited by in corpus (36)
- Diffusion Models in Vision: A Survey
- Inverse-design of nonlinear mechanical metamaterials via video denoising diffusion models
- A Physics-informed Diffusion Model for High-fidelity Flow Field Reconstruction
- How to DP-fy ML: A Practical Guide to Machine Learning with Differential Privacy
- DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models
- Conditional Diffusion Models for Semantic 3D Brain MRI Synthesis
- Diffusion Models, Image Super-Resolution And Everything: A Survey
- Diffusion Model-Based Image Editing: A Survey
- LDMVFI: Video Frame Interpolation with Latent Diffusion Models
- Synthetically Enhanced: Unveiling Synthetic Data's Potential in Medical Imaging Research
- Creativity and Machine Learning: A Survey
- Vector Quantized Diffusion Model for Text-to-Image Synthesis
- MatFusion: A Generative Diffusion Model for SVBRDF Capture
- Dataset Regeneration for Sequential Recommendation
- DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
- Latent Diffusion Model for Conditional Reservoir Facies Generation
- DiffDance: Cascaded Human Motion Diffusion Model for Dance Generation
- 360-Degree Panorama Generation from Few Unregistered NFoV Images
- Explaining generative diffusion models via visual analysis for interpretable decision-making process
- HanDiffuser: Text-to-Image Generation With Realistic Hand Appearances
- 4D Facial Expression Diffusion Model
- ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation
- SegDiff: Image Segmentation with Diffusion Probabilistic Models
- 3D Brain and Heart Volume Generative Models: A Survey
- Geometric-Facilitated Denoising Diffusion Model for 3D Molecule Generation
- Prompt Optimizer of Text-to-Image Diffusion Models for Abstract Concept Understanding
- Toward Scalable Image Feature Compression: A Content-Adaptive and Diffusion-Based Approach
- Art Notions in the Age of (Mis)anthropic AI
- Unleashing Transformers: Parallel Token Prediction with Discrete Absorbing Diffusion for Fast High-Resolution Image Generation from Vector-Quantized Codes
- RGB-D-Fusion: Image Conditioned Depth Diffusion of Humanoid Subjects
- Learning multi-scale local conditional probability models of images
- GSEditPro: 3D Gaussian Splatting Editing with Attention-based Progressive Localization
- Stochastic Super-resolution of Cosmological Simulations with Denoising Diffusion Models
- Space-scale Exploration of the Poor Reliability of Deep Learning Models: the Case of the Remote Sensing of Rooftop Photovoltaic Systems
- Conditionally Strongly Log-Concave Generative Models
- Adapt and Diffuse: Sample-adaptive Reconstruction via Latent Diffusion Models