Hierarchical Text-Conditional Image Generation with CLIP Latents
arXiv:2204.06125
Abstract
Contrastive models like CLIP have been shown to learn robust representations of images that capture both semantics and style. To leverage these representations for image generation, we propose a two-stage model: a prior that generates a CLIP image embedding given a text caption, and a decoder that generates an image conditioned on the image embedding. We show that explicitly generating image representations improves image diversity with minimal loss in photorealism and caption similarity. Our decoders conditioned on image representations can also produce variations of an image that preserve both its semantics and style, while varying the non-essential details absent from the image representation. Moreover, the joint embedding space of CLIP enables language-guided image manipulations in a zero-shot fashion. We use diffusion models for the decoder and experiment with both autoregressive and diffusion models for the prior, finding that the latter are computationally more efficient and produce higher-quality samples.
Cited by in corpus (32)
- The Programmer's Assistant: Conversational Interaction with a Large Language Model for Software Development
- Automated data processing and feature engineering for deep learning and big data applications: a survey
- PromptMagician: Interactive Prompt Engineering for Text-to-Image Creation
- Machine Culture
- RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise Expressions
- DiffEdit: Diffusion-based semantic image editing with mask guidance
- PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like Interactions
- Denoising diffusion algorithm for inverse design of microstructures with fine-tuned nonlinear material properties
- Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance Flow
- Choice Over Control: How Users Write with Large Language Models using Diegetic and Non-Diegetic Prompting
- C2Ideas: Supporting Creative Interior Color Design Ideation with Large Language Model
- A Survey of AI Text-to-Image and AI Text-to-Video Generators
- RAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!
- AVscript: Accessible Video Editing with Audio-Visual Scripts
- PhotoMat: A Material Generator Learned from Single Flash Photos
- PADL: Language-Directed Physics-Based Character Control
- From Paper to Card: Transforming Design Implications with Generative AI
- Who Is Alyx? A new Behavioral Biometric Dataset for User Identification in XR
- CgT-GAN: CLIP-guided Text GAN for Image Captioning
- Latent Denoising Diffusion GAN: Faster sampling, Higher image quality
- Iterative -(de)Blending: a Minimalist Deterministic Diffusion Model
- MOSA: Music Motion with Semantic Annotation Dataset for Cross-Modal Music Processing
- RGB-D-Fusion: Image Conditioned Depth Diffusion of Humanoid Subjects
- Learning multi-scale local conditional probability models of images
- Face Aging via Diffusion-based Editing
- Benchmarking the Fairness of Image Upsampling Methods
- Concept Lens: Visually Analyzing the Consistency of Semantic Manipulation in GANs
- Diffeomorphic Transformations for Time Series Analysis: An Efficient Approach to Nonlinear Warping
- On the Potential of CLIP for Compositional Logical Reasoning
- Conditionally Strongly Log-Concave Generative Models
- Debiasing Sentence Embedders through Contrastive Word Pairs
- Semantic Generative Augmentations for Few-Shot Counting