Vector-quantized Image Modeling with Improved VQGAN
arXiv:2110.04627
Abstract
Pretraining language models with next-token prediction on massive text corpora has delivered phenomenal zero-shot, few-shot, transfer learning and multi-tasking capabilities on both generative and discriminative language tasks. Motivated by this success, we explore a Vector-quantized Image Modeling (VIM) approach that involves pretraining a Transformer to predict rasterized image tokens autoregressively. The discrete image tokens are encoded from a learned Vision-Transformer-based VQGAN (ViT-VQGAN). We first propose multiple improvements over vanilla VQGAN from architecture to codebook learning, yielding better efficiency and reconstruction fidelity. The improved ViT-VQGAN further improves vector-quantized image modeling tasks, including unconditional, class-conditioned image generation and unsupervised representation learning. When trained on ImageNet at \(256\times256\) resolution, we achieve Inception Score (IS) of 175.1 and Fr'echet Inception Distance (FID) of 4.17, a dramatic improvement over the vanilla VQGAN, which obtains 70.6 and 17.04 for IS and FID, respectively. Based on ViT-VQGAN and unsupervised pretraining, we further evaluate the pretrained Transformer by averaging intermediate features, similar to Image GPT (iGPT). This ImageNet-pretrained VIM-L significantly beats iGPT-L on linear-probe accuracy from 60.3% to 73.2% for a similar model size. VIM-L also outperforms iGPT-XL which is trained with extra web image data and larger model size.
Accepted in ICLR 2022
References in corpus (23)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Auto-Encoding Variational Bayes
- Decoupled Weight Decay Regularization
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
- Large Scale GAN Training for High Fidelity Natural Image Synthesis
- Bootstrap your own latent: A new approach to self-supervised Learning
- Language Models are Few-Shot Learners
- Neural Discrete Representation Learning
- Self-Attention Generative Adversarial Networks
- Diffusion Models Beat GANs on Image Synthesis
- Improved Baselines with Momentum Contrastive Learning
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Zero-Shot Text-to-Image Generation
- Generative Modeling by Estimating Gradients of the Data Distribution
- BEiT: BERT Pre-Training of Image Transformers
- Semi-supervised Sequence Learning
- Big Self-Supervised Models are Strong Semi-Supervised Learners
- Large Scale Adversarial Representation Learning
- NVAE: A Deep Hierarchical Variational Autoencoder
- Image Representations Learned With Unsupervised Pre-Training Contain Human-like Biases
- Generating Diverse High-Fidelity Images with VQ-VAE-2
- A Note on Data Biases in Generative Models