Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
arXiv:2206.10789
Abstract
We present the Pathways Autoregressive Text-to-Image (Parti) model, which generates high-fidelity photorealistic images and supports content-rich synthesis involving complex compositions and world knowledge. Parti treats text-to-image generation as a sequence-to-sequence modeling problem, akin to machine translation, with sequences of image tokens as the target outputs rather than text tokens in another language. This strategy can naturally tap into the rich body of prior work on large language models, which have seen continued advances in capabilities and performance through scaling data and model sizes. Our approach is simple: First, Parti uses a Transformer-based image tokenizer, ViT-VQGAN, to encode images as sequences of discrete tokens. Second, we achieve consistent quality improvements by scaling the encoder-decoder Transformer model up to 20B parameters, with a new state-of-the-art zero-shot FID score of 7.23 and finetuned FID score of 3.22 on MS-COCO. Our detailed analysis on Localized Narratives as well as PartiPrompts (P2), a new holistic benchmark of over 1600 English prompts, demonstrate the effectiveness of Parti across a wide variety of categories and difficulty aspects. We also explore and highlight limitations of our models in order to define and exemplify key areas of focus for further improvements. See https://parti.research.google/ for high-resolution images.
Preprint
Cited by in corpus (20)
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Worldwide AI Ethics: a review of 200 guidelines and recommendations for AI governance
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- SpaText: Spatio-Textual Representation for Controllable Image Generation
- Automated data processing and feature engineering for deep learning and big data applications: a survey
- AI's Regimes of Representation: A Community-centered Study of Text-to-Image Models in South Asia
- A newcomer's guide to deep learning for inverse design in nano-photonics
- Generative AI and Process Systems Engineering: The Next Frontier
- A Prompt Log Analysis of Text-to-Image Generation Systems
- Leveraging Large Language Models for Patient Engagement: The Power of Conversational AI in Digital Health
- Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
- DALLE-URBAN: Capturing the urban design expertise of large text to image transformers
- Opportunities and Challenges of Generative-AI in Finance
- Does CLIP Know My Face?
- Can Artificial Intelligence Reconstruct Ancient Mosaics?
- StereoDiffusion: Training-Free Stereo Image Generation Using Latent Diffusion Models
- Parallel Synthesis for Autoregressive Speech Generation
- GSEditPro: 3D Gaussian Splatting Editing with Attention-based Progressive Localization
- GEMRec: Towards Generative Model Recommendation
- Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities