Structured Denoising Diffusion Models in Discrete State-Spaces
arXiv:2107.03006
Abstract
Denoising diffusion probabilistic models (DDPMs) (Ho et al. 2020) have shown impressive results on image and waveform generation in continuous state spaces. Here, we introduce Discrete Denoising Diffusion Probabilistic Models (D3PMs), diffusion-like generative models for discrete data that generalize the multinomial diffusion model of Hoogeboom et al. 2021, by going beyond corruption processes with uniform transition probabilities. This includes corruption with transition matrices that mimic Gaussian kernels in continuous space, matrices based on nearest neighbors in embedding space, and matrices that introduce absorbing states. The third allows us to draw a connection between diffusion models and autoregressive and mask-based generative models. We show that the choice of transition matrix is an important design decision that leads to improved results in image and text domains. We also introduce a new loss function that combines the variational lower bound with an auxiliary cross entropy loss. For text, this model class achieves strong results on character-level text generation while scaling to large vocabularies on LM1B. On the image dataset CIFAR-10, our models approach the sample quality and exceed the log-likelihood of the continuous-space DDPM model.
10 pages plus references and appendices. First two authors contributed equally
References in corpus (10)
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Score-Based Generative Modeling through Stochastic Differential Equations
- Generating Long Sequences with Sparse Transformers
- Improved Denoising Diffusion Probabilistic Models
- BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model
- Discrete Flows: Invertible Generative Models of Discrete Data
- Argmax Flows and Multinomial Diffusion: Learning Categorical Distributions
- Permutation Invariant Graph Generation via Score-Based Generative Modeling
- Symbolic Music Generation with Diffusion Models
- Insertion-Deletion Transformer
Cited by in corpus (11)
- Spatio-temporal Diffusion Point Processes
- Vector Quantized Diffusion Model for Text-to-Image Synthesis
- How Much is Enough? A Study on Diffusion Times in Score-based Generative Models
- 4D Facial Expression Diffusion Model
- SegDiff: Image Segmentation with Diffusion Probabilistic Models
- TDNetGen: Empowering Complex Network Resilience Prediction with Generative Augmentation of Topology and Dynamics
- Denoising Diffusion Gamma Models
- Unleashing Transformers: Parallel Token Prediction with Discrete Absorbing Diffusion for Fast High-Resolution Image Generation from Vector-Quantized Codes
- Autoregressive Diffusion Models
- Beyond In-Place Corruption: Insertion and Deletion In Denoising Probabilistic Models
- Zero-Shot Translation using Diffusion Models