Latent Denoising Diffusion GAN: Faster sampling, Higher image quality
arXiv:2406.11713 · doi:10.1109/ACCESS.2024.3406535
Abstract
Diffusion models are emerging as powerful solutions for generating high-fidelity and diverse images, often surpassing GANs under many circumstances. However, their slow inference speed hinders their potential for real-time applications. To address this, DiffusionGAN leveraged a conditional GAN to drastically reduce the denoising steps and speed up inference. Its advancement, Wavelet Diffusion, further accelerated the process by converting data into wavelet space, thus enhancing efficiency. Nonetheless, these models still fall short of GANs in terms of speed and image quality. To bridge these gaps, this paper introduces the Latent Denoising Diffusion GAN, which employs pre-trained autoencoders to compress images into a compact latent space, significantly improving inference speed and image quality. Furthermore, we propose a Weighted Learning strategy to enhance diversity and image quality. Experimental results on the CIFAR-10, CelebA-HQ, and LSUN-Church datasets prove that our model achieves state-of-the-art running speed among diffusion models. Compared to its predecessors, DiffusionGAN and Wavelet Diffusion, our model shows remarkable improvements in all evaluation metrics. Code and pre-trained checkpoints: \url{https://github.com/thanhluantrinh/LDDGAN.git}
Submited to IEEE Access
References in corpus (22)
- Generative Adversarial Networks
- Spectral Normalization for Generative Adversarial Networks
- Stochastic Backpropagation and Approximate Inference in Deep Generative Models
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Diffusion Models Beat GANs on Image Synthesis
- NICE: Non-linear Independent Components Estimation
- LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop
- Generative Modeling by Estimating Gradients of the Data Distribution
- Training Generative Adversarial Networks with Limited Data
- Which Training Methods for GANs do actually Converge?
- Imagen Video: High Definition Video Generation with Diffusion Models
- Blended Latent Diffusion
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
- Improved Precision and Recall Metric for Assessing Generative Models
- Vector-quantized Image Modeling with Improved VQGAN
- Analytic-DPM: an Analytic Estimate of the Optimal Reverse Variance in Diffusion Probabilistic Models
- Notes on Kullback-Leibler Divergence and Likelihood
- EGSDE: Unpaired Image-to-Image Translation via Energy-Guided Stochastic Differential Equations
- Knowledge Distillation in Iterative Generative Models for Improved Sampling Speed
- On Fast Sampling of Diffusion Probabilistic Models
- Noise Estimation for Generative Diffusion Models