PromptPaint: Steering Text-to-Image Generation Through Paint Medium-like Interactions
arXiv:2308.05184 · doi:10.1145/3586183.3606777
Abstract
While diffusion-based text-to-image (T2I) models provide a simple and powerful way to generate images, guiding this generation remains a challenge. For concepts that are difficult to describe through language, users may struggle to create prompts. Moreover, many of these models are built as end-to-end systems, lacking support for iterative shaping of the image. In response, we introduce PromptPaint, which combines T2I generation with interactions that model how we use colored paints. PromptPaint allows users to go beyond language to mix prompts that express challenging concepts. Just as we iteratively tune colors through layered placements of paint on a physical canvas, PromptPaint similarly allows users to apply different prompts to different canvas areas and times of the generative process. Through a set of studies, we characterize different approaches for mixing prompts, design trade-offs, and socio-technical challenges for generative models. With PromptPaint we provide insight into future steerable generative tools.
Accepted to UIST2023
References in corpus (19)
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Diffusion Models Beat GANs on Image Synthesis
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Evaluating Large Language Models Trained on Code
- A Learned Representation For Artistic Style
- Classifier-Free Diffusion Guidance
- DreamFusion: Text-to-3D using 2D Diffusion
- Prompt-to-Prompt Image Editing with Cross Attention Control
- Make-A-Video: Text-to-Video Generation without Text-Video Data
- MusicLM: Generating Music From Text
- RePrompt: Automatic Prompt Editing to Refine AI-Generative Art Towards Precise Expressions
- Cold Diffusion: Inverting Arbitrary Image Transforms Without Noise
- DiffEdit: Diffusion-based semantic image editing with mask guidance
- Sequential Gallery for Interactive Visual Design Optimization
- Phenaki: Variable Length Video Generation From Open Domain Textual Description
- Composer: Creative and Controllable Image Synthesis with Composable Conditions
- GANSlider: How Users Control Generative Models for Images using Multiple Sliders with and without Feedforward Information
- MagicMix: Semantic Mixing with Diffusion Models
- Zero-shot Image-to-Image Translation