GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
arXiv:2112.10741
Abstract
Diffusion models have recently been shown to generate high-quality synthetic images, especially when paired with a guidance technique to trade off diversity for fidelity. We explore diffusion models for the problem of text-conditional image synthesis and compare two different guidance strategies: CLIP guidance and classifier-free guidance. We find that the latter is preferred by human evaluators for both photorealism and caption similarity, and often produces photorealistic samples. Samples from a 3.5 billion parameter text-conditional diffusion model using classifier-free guidance are favored by human evaluators to those from DALL-E, even when the latter uses expensive CLIP reranking. Additionally, we find that our models can be fine-tuned to perform image inpainting, enabling powerful text-driven image editing. We train a smaller model on a filtered dataset and release the code and weights at https://github.com/openai/glide-text2im.
20 pages, 18 figures
Cited by in corpus (51)
- Diffusion Models in Vision: A Survey
- Blended Diffusion for Text-driven Editing of Natural Images
- Towards Small Object Editing: A Benchmark Dataset and A Training-Free Approach
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- SpaText: Spatio-Textual Representation for Controllable Image Generation
- Automated data processing and feature engineering for deep learning and big data applications: a survey
- DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models
- AI Art in Architecture
- RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise Expressions
- DiffEdit: Diffusion-based semantic image editing with mask guidance
- A survey of recent methods for addressing AI fairness and bias in biomedicine
- PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like Interactions
- Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance Flow
- Choice Over Control: How Users Write with Large Language Models using Diegetic and Non-Diegetic Prompting
- What Does DALL-E 2 Know About Radiology?
- AVscript: Accessible Video Editing with Audio-Visual Scripts
- SGDiff: A Style Guided Diffusion Model for Fashion Synthesis
- Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
- DALLE-URBAN: Capturing the urban design expertise of large text to image transformers
- PADL: Language-Directed Physics-Based Character Control
- DiffDance: Cascaded Human Motion Diffusion Model for Dance Generation
- A Text-guided Protein Design Framework
- Sparse Visual Counterfactual Explanations in Image Space
- Controllable Data Generation by Deep Learning: A Review
- CHeart: A Conditional Spatio-Temporal Generative Model for Cardiac Anatomy
- Painterly Image Harmonization using Diffusion Model
- SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image Models
- BenthicNet: A global compilation of seafloor images for deep learning applications
- Does CLIP Know My Face?
- ElasticDiffusion: Training-free Arbitrary Size Image Generation through Global-Local Content Separation
- Probing the Limits and Capabilities of Diffusion Models for the Anatomic Editing of Digital Twins
- Perturbing Attention Gives You More Bang for the Buck: Subtle Imaging Perturbations That Efficiently Fool Customized Diffusion Models
- Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation
- PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
- Multi-modal Machine Learning for Vehicle Rating Predictions Using Image, Text, and Parametric Data
- A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
- Translation-Enhanced Multilingual Text-to-Image Generation
- LEDITS: Real Image Editing with DDPM Inversion and Semantic Guidance
- Interactive Neural Painting
- Structured Generative Models for Scene Understanding
- EBDM: Exemplar-guided Image Translation with Brownian-bridge Diffusion Models
- Steganography Beyond Space-Time with Chain of Multimodal AI
- Comprehensive Dataset of Synthetic and Manipulated Overhead Imagery for Development and Evaluation of Forensic Tools
- RGB-D-Fusion: Image Conditioned Depth Diffusion of Humanoid Subjects
- Face Aging via Diffusion-based Editing
- SketchBetween: Video-to-Video Synthesis for Sprite Animation via Sketches
- Image Captions are Natural Prompts for Text-to-Image Models
- Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities
- Adapt and Diffuse: Sample-adaptive Reconstruction via Latent Diffusion Models
- CRAFT: Cultural Russian-Oriented Dataset Adaptation for Focused Text-to-Image Generation
- Semantic Generative Augmentations for Few-Shot Counting