Generating Diverse High-Fidelity Images with VQ-VAE-2
arXiv:1906.00446
Abstract
We explore the use of Vector Quantized Variational AutoEncoder (VQ-VAE) models for large scale image generation. To this end, we scale and enhance the autoregressive priors used in VQ-VAE to generate synthetic samples of much higher coherence and fidelity than possible before. We use simple feed-forward encoder and decoder networks, making our model an attractive candidate for applications where the encoding and/or decoding speed is critical. Additionally, VQ-VAE requires sampling an autoregressive model only in the compressed latent space, which is an order of magnitude faster than sampling in the pixel space, especially for large images. We demonstrate that a multi-scale hierarchical organization of VQ-VAE, augmented with powerful priors over the latent codes, is able to generate samples with quality that rivals that of state of the art Generative Adversarial Networks on multifaceted datasets such as ImageNet, while not suffering from GAN's known shortcomings such as mode collapse and lack of diversity.
References in corpus (5)
Cited by in corpus (21)
- Diffusion Models Beat GANs on Image Synthesis
- Zero-Shot Text-to-Image Generation
- Controlling generative models with continuous factors of variations
- Cross-Camera Feature Prediction for Intra-Camera Supervised Person Re-identification across Distant Scenes
- On Training Sample Memorization: Lessons from Benchmarking Generative Modeling with a Large-scale Competition
- A unified framework for 21cm tomography sample generation and parameter inference with Progressively Growing GANs
- Non Gaussian Denoising Diffusion Models
- A Generative Model of Galactic Dust Emission Using Variational Inference
- VQ-GNN: A Universal Framework to Scale up Graph Neural Networks using Vector Quantization
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech
- Vector Quantized Contrastive Predictive Coding for Template-based Music Generation
- Bilateral Denoising Diffusion Models
- High- and Low-level image component decomposition using VAEs for improved reconstruction and anomaly detection
- The Multi-speaker Multi-style Voice Cloning Challenge 2021
- Augmentation-Interpolative AutoEncoders for Unsupervised Few-Shot Image Generation
- VQ-DRAW: A Sequential Discrete VAE
- Set Distribution Networks: a Generative Model for Sets of Images
- D2C: Diffusion-Denoising Models for Few-shot Conditional Generation
- Zero-Shot Translation using Diffusion Models
- IB-DRR: Incremental Learning with Information-Back Discrete Representation Replay
- NP-DRAW: A Non-Parametric Structured Latent Variable Model for Image Generation