Diffusion Models in Vision: A Survey
arXiv:2209.04747 · doi:10.1109/TPAMI.2023.3261988
Abstract
Denoising diffusion models represent a recent emerging topic in computer vision, demonstrating remarkable results in the area of generative modeling. A diffusion model is a deep generative model that is based on two stages, a forward diffusion stage and a reverse diffusion stage. In the forward diffusion stage, the input data is gradually perturbed over several steps by adding Gaussian noise. In the reverse stage, a model is tasked at recovering the original input data by learning to gradually reverse the diffusion process, step by step. Diffusion models are widely appreciated for the quality and diversity of the generated samples, despite their known computational burdens, i.e. low speeds due to the high number of steps involved during sampling. In this survey, we provide a comprehensive review of articles on denoising diffusion models applied in vision, comprising both theoretical and practical contributions in the field. First, we identify and present three generic diffusion modeling frameworks, which are based on denoising diffusion probabilistic models, noise conditioned score networks, and stochastic differential equations. We further discuss the relations between diffusion models and other deep generative models, including variational auto-encoders, generative adversarial networks, energy-based models, autoregressive models and normalizing flows. Then, we introduce a multi-perspective categorization of diffusion models applied in computer vision. Finally, we illustrate the current limitations of diffusion models and envision some interesting directions for future research.
Accepted in IEEE Transactions on Pattern Analysis and Machine Intelligence. 25 pages, 3 figures
References in corpus (48)
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Diffusion Models Beat GANs on Image Synthesis
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Temporal Ensembling for Semi-Supervised Learning
- NICE: Non-linear Independent Components Estimation
- Score-Based Generative Modeling through Stochastic Differential Equations
- Classifier-Free Diffusion Guidance
- Blended Latent Diffusion
- Elucidating the Design Space of Diffusion-Based Generative Models
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps
- Amortised MAP Inference for Image Super-resolution
- Tackling the Generative Learning Trilemma with Denoising Diffusion GANs
- Flexible Diffusion Modeling of Long Videos
- Pretraining is All You Need for Image-to-Image Translation
- Semantic Image Synthesis via Diffusion Models
- Diffusion Models for Adversarial Purification
- Analytic-DPM: an Analytic Estimate of the Optimal Reverse Variance in Diffusion Probabilistic Models
- Diffusion Models for Implicit Image Segmentation Ensembles
- UNIT-DDPM: UNpaired Image Translation with Denoising Diffusion Probabilistic Models
- EGSDE: Unpaired Image-to-Image Translation via Energy-Guided Stochastic Differential Equations
- Sliced Score Matching: A Scalable Approach to Density and Score Estimation
- Diffusion-GAN: Training GANs with Diffusion
- Fast Sampling of Diffusion Models with Exponential Integrator
- On Fast Sampling of Diffusion Probabilistic Models
- ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis
- Diffusion Models for Video Prediction and Infilling
- Learning to Efficiently Sample from Diffusion Probabilistic Models
- A Variational Perspective on Diffusion-Based Generative Models and Score Matching
- Accelerating Diffusion Models via Early Stop of the Diffusion Process
- Text-Guided Synthesis of Artistic Images with Retrieval-Augmented Diffusion Models
- Subspace Diffusion Generative Models
- Unsupervised Medical Image Translation with Adversarial Diffusion Models
- Learning Fast Samplers for Diffusion Models by Differentiating Through Sample Quality
- Few-Shot Diffusion Models
- Gotta Go Fast When Generating Data with Score-Based Models
- The Swiss Army Knife for Image-to-Image Translation: Multi-Task Diffusion Models
- Non Gaussian Denoising Diffusion Models
- Restoring Vision in Adverse Weather Conditions with Patch-Based Denoising Diffusion Models
- Diffusion Causal Models for Counterfactual Estimation
- DiVAE: Photorealistic Images Synthesis with Denoising Diffusion Decoder
- Bilateral Denoising Diffusion Models
- Maximum Likelihood Training of Implicit Nonlinear Diffusion Models
- On Analyzing Generative and Denoising Capabilities of Diffusion-based Deep Generative Models
- Denoising Likelihood Score Matching for Conditional Score-based Data Generation
- On Conditioning the Input Noise for Controlled Image Generation with Diffusion Models
- Non-Uniform Diffusion Models
- Accelerating Score-based Generative Models for High-Resolution Image Synthesis
- Heavy-tailed denoising score matching
Cited by in corpus (76)
- Explainable Artificial Intelligence (XAI) 2.0: A Manifesto of Open Challenges and Interdisciplinary Research Directions
- Ten Years of Generative Adversarial Nets (GANs): A survey of the state-of-the-art
- Generative AI in the Construction Industry: Opportunities & Challenges
- SpaText: Spatio-Textual Representation for Controllable Image Generation
- Automated data processing and feature engineering for deep learning and big data applications: a survey
- A survey on GANs for computer vision: Recent research, analysis and taxonomy
- Diffusion Model-Based Image Editing: A Survey
- SpectralDiff: A Generative Framework for Hyperspectral Image Classification with Diffusion Models
- Adaptive Latent Diffusion Model for 3D Medical Image to Image Translation: Multi-modal Magnetic Resonance Imaging Study
- Deep Learning for Retrospective Motion Correction in MRI: A Comprehensive Review
- A newcomer's guide to deep learning for inverse design in nano-photonics
- Diffusion Models Meet Remote Sensing: Principles, Methods, and Perspectives
- Diff-MTS: Temporal-Augmented Conditional Diffusion-based AIGC for Industrial Time Series Towards the Large Model Era
- Development of Skip Connection in Deep Neural Networks for Computer Vision and Medical Image Analysis: A Survey
- FDiff-Fusion:Denoising diffusion fusion network based on fuzzy learning for 3D medical image segmentation
- FakeNews: GAN-based generation of realistic 3D volumetric data -- A systematic review and taxonomy
- Artificial Intelligence and Deep Learning Algorithms for Epigenetic Sequence Analysis: A Review for Epigeneticists and AI Experts
- FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution
- Quantum-Noise-Driven Generative Diffusion Models
- Assessing and Understanding Creativity in Large Language Models
- On the design space between molecular mechanics and machine learning force fields
- Diffusion Models as Stochastic Quantization in Lattice Field Theory
- Creating synthetic energy meter data using conditional diffusion and building metadata
- RNDiff: Rainfall nowcasting with Condition Diffusion Model
- A Review of Emerging Research Directions in Abstract Visual Reasoning
- Reconstruction from edge image combined with color and gradient difference for industrial surface anomaly detection
- Conditional score-based diffusion models for solving inverse problems in mechanics
- FAIR: Frequency-aware Image Restoration for Industrial Visual Anomaly Detection
- Robust deep learning for eye fundus images: Bridging real and synthetic data for enhancing generalization
- Controllable Generation with Text-to-Image Diffusion Models: A Survey
- Face Morphing Attack Detection with Denoising Diffusion Probabilistic Models
- An Integrated Fusion Framework for Ensemble Learning Leveraging Gradient Boosting and Fuzzy Rule-Based Models
- A Shift In Artistic Practices through Artificial Intelligence
- Transfer learning with generative models for object detection on limited datasets
- The stochastic digital human is now enrolling for in silico imaging trials -- Methods and tools for generating digital cohorts
- Machine Learning in Acoustics: A Review and Open-Source Repository
- A Multi-scale Generalized Shrinkage Threshold Network for Image Blind Deblurring in Remote Sensing
- Survey on Visual Signal Coding and Processing with Generative Models: Technologies, Standards and Optimization
- GenSelfDiff-HIS: Generative Self-Supervision Using Diffusion for Histopathological Image Segmentation
- A Survey on Personalized Content Synthesis with Diffusion Models
- Physics-informed generative model for drug-like molecule conformers
- SeisFusion: Constrained Diffusion Model with Input Guidance for 3D Seismic Data Interpolation and Reconstruction
- Leave No One Behind: Fairness-Aware Cross-Domain Recommender Systems for Non-Overlapping Users
- 3D Multiphase Heterogeneous Microstructure Generation Using Conditional Latent Diffusion Models
- Synthetic Simplicity: Unveiling Bias in Medical Data Augmentation
- When Generative Artificial Intelligence meets Extended Reality: A Systematic Review
- Towards Effective Data-Free Knowledge Distillation via Diverse Diffusion Augmentation
- Ultra fast, event-by-event heavy-ion simulations for next generation experiments
- Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval
- CleAR: Robust Context-Guided Generative Lighting Estimation for Mobile Augmented Reality
- Towards Natural Machine Unlearning
- Controllable Face Synthesis with Semantic Latent Diffusion Models
- A Survey on Human Interaction Motion Generation
- Generative Design of a Gas Turbine Combustor Using Invertible Neural Networks
- Denoising Diffusion Probabilistic Model for realistic and fast generated \textit{Euclid}-like data for weak lensing analysis
- Toward a foundation model for heavy-ion collision experiments based on point-cloud diffusion
- A Survey on Text-Driven 360-Degree Panorama Generation
- Conditional diffusion model for inverse prediction of process parameters and dendritic microstructures from mechanical properties
- Vision-Language Modeling with Regularized Spatial Transformer Networks for All Weather Crosswind Landing of Aircraft
- AI-rays: Exploring Bias in the Gaze of AI Through a Multimodal Interactive Installation
- Neural Network-Based Bandit: A Medium Access Control for the IIoT Alarm Scenario
- A Deep-Learning Framework for Land-Sliding Classification from Remote Sensing Image
- Stochastic Schrödinger Equations for Quantum Reverse Diffusion
- Diff-GO: Enhancing Diffusion Models for Goal-Oriented Communications
- Predicting the Dynamics of Complex System via Multiscale Diffusion Autoencoder
- Diffusion-Driven Inertial Generated Data for Smartphone Location Classification
- RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
- Beyond Consistency: Preserving Temporal Structure in Zero-Shot Video Editing
- DAWN-FM: Data-Aware and Noise-Informed Flow Matching for Solving Inverse Problems
- Diffusion model for relational inference
- Group-Equivariant Diffusion Models for Lattice Field Theory
- Hyperspectral data augmentation with transformer-based diffusion models
- A Differentiable Surrogate Model for the Generation of Radio Pulses from In-Ice Neutrino Interactions
- A Diffusion-based Data Generator for Training Object Recognition Models in Ultra-Range Distance
- SinFormer: A Tailored Transformer for Robust Radio Frequency Fingerprint Identification
- DeepSeek-Inspired Exploration of RL-based LLMs and Synergy with Wireless Networks: A Survey