SpaText: Spatio-Textual Representation for Controllable Image Generation
arXiv:2211.14305 · doi:10.1109/CVPR52729.2023.01762
Abstract
Recent text-to-image diffusion models are able to generate convincing results of unprecedented quality. However, it is nearly impossible to control the shapes of different regions/objects or their layout in a fine-grained fashion. Previous attempts to provide such controls were hindered by their reliance on a fixed set of labels. To this end, we present SpaText - a new method for text-to-image generation using open-vocabulary scene control. In addition to a global text prompt that describes the entire scene, the user provides a segmentation map where each region of interest is annotated by a free-form natural language description. Due to lack of large-scale datasets that have a detailed textual description for each region in the image, we choose to leverage the current large-scale text-to-image datasets and base our approach on a novel CLIP-based spatio-textual representation, and show its effectiveness on two state-of-the-art diffusion models: pixel-based and latent-based. In addition, we show how to extend the classifier-free guidance method in diffusion models to the multi-conditional case and present an alternative accelerated inference algorithm. Finally, we offer several automatic evaluation metrics and use them, in addition to FID scores and a user study, to evaluate our method and show that it achieves state-of-the-art results on image generation with free-form textual scene control.
CVPR 2023. Project page available at: https://omriavrahami.com/spatext
References in corpus (31)
- Auto-Encoding Variational Bayes
- Explaining and Harnessing Adversarial Examples
- Denoising Diffusion Probabilistic Models
- Intriguing properties of neural networks
- Neural Discrete Representation Learning
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Generative Adversarial Text to Image Synthesis
- Diffusion Models Beat GANs on Image Synthesis
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Diffusion Models in Vision: A Survey
- Deep Unsupervised Learning using Nonequilibrium Thermodynamics
- Zero-Shot Text-to-Image Generation
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
- Blended Diffusion for Text-driven Editing of Natural Images
- Classifier-Free Diffusion Guidance
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
- CogView: Mastering Text-to-Image Generation via Transformers
- Prompt-to-Prompt Image Editing with Cross Attention Control
- Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
- Blended Latent Diffusion
- Semantic Object Accuracy for Generative Text-to-Image Synthesis
- AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks
- CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers
- DiffEdit: Diffusion-based semantic image editing with mask guidance
- Diffusion Models already have a Semantic Latent Space
- KNN-Diffusion: Image Generation via Large-Scale Retrieval
- Re-Imagen: Retrieval-Augmented Text-to-Image Generator
- Text-Guided Synthesis of Artistic Images with Retrieval-Augmented Diffusion Models
- Text2LIVE: Text-Driven Layered Image and Video Editing
- Controlling Style and Semantics in Weakly-Supervised Image Generation
- High-Resolution Image Editing via Multi-Stage Blended Diffusion
Cited by in corpus (10)
- Break-A-Scene: Extracting Multiple Concepts from a Single Image
- Diffusion Model-Based Image Editing: A Survey
- The Chosen One: Consistent Characters in Text-to-Image Diffusion Models
- Controllable Generation with Text-to-Image Diffusion Models: A Survey
- Stable Flow: Vital Layers for Training-Free Image Editing
- DiffUHaul: A Training-Free Method for Object Dragging in Images
- LAPIG: Language Guided Projector Image Generation with Surface Adaptation and Stylization
- ReCorD: Reasoning and Correcting Diffusion for HOI Generation
- When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following Ability
- Pro-DG: Procedural Diffusion Guidance for Architectural Facade Generation