Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images
arXiv:2011.10650
Abstract
We present a hierarchical VAE that, for the first time, generates samples quickly while outperforming the PixelCNN in log-likelihood on all natural image benchmarks. We begin by observing that, in theory, VAEs can actually represent autoregressive models, as well as faster, better models if they exist, when made sufficiently deep. Despite this, autoregressive models have historically outperformed VAEs in log-likelihood. We test if insufficient depth explains why by scaling a VAE to greater stochastic depth than previously explored and evaluating it CIFAR-10, ImageNet, and FFHQ. In comparison to the PixelCNN, these very deep VAEs achieve higher likelihoods, use fewer parameters, generate samples thousands of times faster, and are more easily applied to high-resolution images. Qualitative studies suggest this is because the VAE learns efficient hierarchical visual representations. We release our source code and models at https://github.com/openai/vdvae.
17 pages, 14 figures
References in corpus (7)
- Language Models are Few-Shot Learners
- NICE: Non-linear Independent Components Estimation
- Generating Long Sequences with Sparse Transformers
- Variational Lossy Autoencoder
- Fixup Initialization: Residual Learning Without Normalization
- Jukebox: A Generative Model for Music
- Learnable Explicit Density for Continuous Latent Space and Variational Inference
Cited by in corpus (9)
- Diffusion Models Beat GANs on Image Synthesis
- Improved Denoising Diffusion Probabilistic Models
- Knowledge Distillation in Iterative Generative Models for Improved Sampling Speed
- Learning to Efficiently Sample from Diffusion Probabilistic Models
- VARA-TTS: Non-Autoregressive Text-to-Speech Synthesis based on Very Deep VAE with Residual Attention
- Clockwork Variational Autoencoders
- Re-parameterizing VAEs for stability
- D2C: Diffusion-Denoising Models for Few-shot Conditional Generation
- Cascading Modular Network (CAM-Net) for Multimodal Image Synthesis