The Creativity of Text-to-Image Generation
arXiv:2206.02904 · doi:10.1145/3569219.3569352
Abstract
Text-guided synthesis of images has made a giant leap towards becoming a mainstream phenomenon. With text-to-image generation systems, anybody can create digital images and artworks. This provokes the question of whether text-to-image generation is creative. This paper expounds on the nature of human creativity involved in text-to-image art (so-called "AI art") with a specific focus on the practice of prompt engineering. The paper argues that the current product-centered view of creativity falls short in the context of text-to-image generation. A case exemplifying this shortcoming is provided and the importance of online communities for the creative ecosystem of text-to-image art is highlighted. The paper provides a high-level summary of this online ecosystem drawing on Rhodes' conceptual four P model of creativity. Challenges for evaluating the creativity of text-to-image generation and opportunities for research on text-to-image generation in the field of Human-Computer Interaction (HCI) are discussed.
References in corpus (4)
Cited by in corpus (22)
- The Creativity of Text-to-Image Generation
- A Taxonomy of Prompt Modifiers for Text-To-Image Generation
- Using Text-to-Image Generation for Architectural Design Ideation
- Machine Culture
- What Does DALL-E 2 Know About Radiology?
- "An Adapt-or-Die Type of Situation": Perception, Adoption, and Use of Text-To-Image-Generation AI by Game Industry Professionals
- Generative AI in the Wild: Prospects, Challenges, and Strategies
- A Prompt Log Analysis of Text-to-Image Generation Systems
- ABScribe: Rapid Exploration & Organization of Multiple Writing Variations in Human-AI Co-Writing Tasks using Large Language Models
- Exploring Perspectives on the Impact of Artificial Intelligence on the Creativity of Knowledge Work: Beyond Mechanised Plagiarism and Stochastic Parrots
- Perceptions and Realities of Text-to-Image Generation
- Domain-specific ChatBots for Science using Embeddings
- The Human-GenAI Value Loop in Human-Centered Innovation: Beyond the Magical Narrative
- Algorithmic Ways of Seeing: Using Object Detection to Facilitate Art Exploration
- DeTiME: Diffusion-Enhanced Topic Modeling using Encoder-decoder based LLM
- Situating the social issues of image generation models in the model life cycle: a sociotechnical approach
- GenColor: Generative Color-Concept Association in Visual Design
- Galaxy Imaging with Generative Models: Insights from a Two-Models Framework
- Unlimited Editions: Documenting Human Style in AI Art Generation
- Investigating the diversity and stylization of contemporary user generated visual arts in the complexity entropy plane
- Steering Large Text-to-Image Model for Abstract Art Synthesis: Preference-based Prompt Optimization and Visualization
- Offline Evaluation of Set-Based Text-to-Image Generation