Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
arXiv:2205.11487
Abstract
We present Imagen, a text-to-image diffusion model with an unprecedented degree of photorealism and a deep level of language understanding. Imagen builds on the power of large transformer language models in understanding text and hinges on the strength of diffusion models in high-fidelity image generation. Our key discovery is that generic large language models (e.g. T5), pretrained on text-only corpora, are surprisingly effective at encoding text for image synthesis: increasing the size of the language model in Imagen boosts both sample fidelity and image-text alignment much more than increasing the size of the image diffusion model. Imagen achieves a new state-of-the-art FID score of 7.27 on the COCO dataset, without ever training on COCO, and human raters find Imagen samples to be on par with the COCO data itself in image-text alignment. To assess text-to-image models in greater depth, we introduce DrawBench, a comprehensive and challenging benchmark for text-to-image models. With DrawBench, we compare Imagen with recent methods including VQ-GAN+CLIP, Latent Diffusion Models, and DALL-E 2, and find that human raters prefer Imagen over other models in side-by-side comparisons, both in terms of sample quality and image-text alignment. See https://imagen.research.google/ for an overview of the results.
Cited by in corpus (102)
- Diffusion Models in Vision: A Survey
- Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better
- Deep Learning Approaches for Data Augmentation in Medical Imaging: A Review
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Artificial-intelligence-based molecular classification of diffuse gliomas using rapid, label-free optical imaging
- SpaText: Spatio-Textual Representation for Controllable Image Generation
- Automated data processing and feature engineering for deep learning and big data applications: a survey
- DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models
- Using Text-to-Image Generation for Architectural Design Ideation
- Break-A-Scene: Extracting Multiple Concepts from a Single Image
- RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise Expressions
- Diffusion Models, Image Super-Resolution And Everything: A Survey
- DiffEdit: Diffusion-based semantic image editing with mask guidance
- Face Generation and Editing with StyleGAN: A Survey
- SpectralDiff: A Generative Framework for Hyperspectral Image Classification with Diffusion Models
- PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like Interactions
- Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance Flow
- AI's Regimes of Representation: A Community-centered Study of Text-to-Image Models in South Asia
- Learning from models beyond fine-tuning
- Uni-paint: A Unified Framework for Multimodal Image Inpainting with Pretrained Diffusion Model
- Txt2Img-MHN: Remote Sensing Image Generation from Text Using Modern Hopfield Networks
- BerDiff: Conditional Bernoulli Diffusion Model for Medical Image Segmentation
- A Prompt Log Analysis of Text-to-Image Generation Systems
- A Survey of AI Text-to-Image and AI Text-to-Video Generators
- Dynamical Regimes of Diffusion Models
- Generative Artificial Intelligence Meets Synthetic Aperture Radar: A Survey
- AVscript: Accessible Video Editing with Audio-Visual Scripts
- Transparent AI Disclosure Obligations: Who, What, When, Where, Why, How
- DiLightNet: Fine-grained Lighting Control for Diffusion-based Image Generation
- Leveraging Large Language Models for Patient Engagement: The Power of Conversational AI in Digital Health
- Improved machine learning algorithm for predicting ground state properties
- Accelerating Material Design with the Generative Toolkit for Scientific Discovery
- Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
- Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
- RealFill: Reference-Driven Generation for Authentic Image Completion
- DALLE-URBAN: Capturing the urban design expertise of large text to image transformers
- Not Only Generative Art: Stable Diffusion for Content-Style Disentanglement in Art Analysis
- Clinical Insights: A Comprehensive Review of Language Models in Medicine
- Advances in machine-learning-based sampling motivated by lattice quantum chromodynamics
- DiffDance: Cascaded Human Motion Diffusion Model for Dance Generation
- EDMP: Ensemble-of-costs-guided Diffusion for Motion Planning
- The Chosen One: Consistent Characters in Text-to-Image Diffusion Models
- Deep Learning for Optical Tweezers
- RNDiff: Rainfall nowcasting with Condition Diffusion Model
- Neural radiance fields in the industrial and robotics domain: applications, research opportunities and use cases
- UGG: Unified Generative Grasping
- Generative Learning of the Solution of Parametric Partial Differential Equations Using Guided Diffusion Models and Virtual Observations
- HanDiffuser: Text-to-Image Generation With Realistic Hand Appearances
- Score-Based Generative Models for PET Image Reconstruction
- AvatarFusion: Zero-shot Generation of Clothing-Decoupled 3D Avatars Using 2D Diffusion
- 4D Facial Expression Diffusion Model
- Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields
- Seven Useful Questions in Density Functional Theory
- Does CLIP Know My Face?
- Boosting GUI Prototyping with Diffusion Models
- ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation
- Iterative -(de)Blending: a Minimalist Deterministic Diffusion Model
- Make-It-4D: Synthesizing a Consistent Long-Term Dynamic Scene Video from a Single Image
- DiffPhase: Generative Diffusion-based STFT Phase Retrieval
- Face Super-Resolution Using Stochastic Differential Equations
- Perturbing Attention Gives You More Bang for the Buck: Subtle Imaging Perturbations That Efficiently Fool Customized Diffusion Models
- Boosting Latent Diffusion with Flow Matching
- Unleashing the Potential of Pre-Trained Diffusion Models for Generalizable Person Re-Identification
- A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
- Surgical Text-to-Image Generation
- LEDITS: Real Image Editing with DDPM Inversion and Semantic Guidance
- Translation-Enhanced Multilingual Text-to-Image Generation
- F3-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video Synthesis
- Prompt Optimizer of Text-to-Image Diffusion Models for Abstract Concept Understanding
- Leveraging Large Language Models for Collective Decision-Making
- Reminding Forgetful Organic Neuromorphic Device Networks
- Steganography Beyond Space-Time with Chain of Multimodal AI
- ZePo: Zero-Shot Portrait Stylization with Faster Sampling
- CLIP4Sketch: Enhancing Sketch to Mugshot Matching through Dataset Augmentation using Diffusion Models
- Comprehensive Dataset of Synthetic and Manipulated Overhead Imagery for Development and Evaluation of Forensic Tools
- Optimal Linear Subspace Search: Learning to Construct Fast and High-Quality Schedulers for Diffusion Models
- Content-Based Search for Deep Generative Models
- RGB-D-Fusion: Image Conditioned Depth Diffusion of Humanoid Subjects
- Training-free Subject-Enhanced Attention Guidance for Compositional Text-to-image Generation
- When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following Ability
- DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models
- GSEditPro: 3D Gaussian Splatting Editing with Attention-based Progressive Localization
- Image-Text Out-Of-Context Detection Using Synthetic Multimodal Misinformation
- Face Aging via Diffusion-based Editing
- Conditional Distribution Modelling for Few-Shot Image Synthesis with Diffusion Models
- Content-aware Tile Generation using Exterior Boundary Inpainting
- Human Inspired Progressive Alignment and Comparative Learning for Grounded Word Acquisition
- Concept Lens: Visually Analyzing the Consistency of Semantic Manipulation in GANs
- BOTH2Hands: Inferring 3D Hands from Both Text Prompts and Body Dynamics
- TPA3D: Triplane Attention for Fast Text-to-3D Generation
- Image Captions are Natural Prompts for Text-to-Image Models
- A Systematic Review of Open Datasets Used in Text-to-Image (T2I) Gen AI Model Safety
- One-shot Unsupervised Domain Adaptation with Personalized Diffusion Models
- ForceGen: End-to-end de novo protein generation based on nonlinear mechanical unfolding responses using a protein language diffusion model
- Steering Large Text-to-Image Model for Abstract Art Synthesis: Preference-based Prompt Optimization and Visualization
- Generative inpainting of incomplete Euclidean distance matrices of trajectories generated by a fractional Brownian motion
- Diffeomorphic Transformations for Time Series Analysis: An Efficient Approach to Nonlinear Warping
- ESCT3D: Efficient and Selectively Controllable Text-Driven 3D Content Generation with Gaussian Splatting
- Adapt and Diffuse: Sample-adaptive Reconstruction via Latent Diffusion Models
- Investigating Conceptual Blending of a Diffusion Model for Improving Nonword-to-Image Generation
- Semantic Generative Augmentations for Few-Shot Counting
- Conditionally Strongly Log-Concave Generative Models