A Taxonomy of Prompt Modifiers for Text-To-Image Generation
arXiv:2204.13988 · doi:10.1080/0144929X.2023.2286532
Abstract
Text-to-image generation has seen an explosion of interest since 2021. Today, beautiful and intriguing digital images and artworks can be synthesized from textual inputs ("prompts") with deep generative models. Online communities around text-to-image generation and AI generated art have quickly emerged. This paper identifies six types of prompt modifiers used by practitioners in the online community based on a 3-month ethnographic study. The novel taxonomy of prompt modifiers provides researchers a conceptual starting point for investigating the practice of text-to-image generation, but may also help practitioners of AI generated art improve their images. We further outline how prompt modifiers are applied in the practice of "prompt engineering." We discuss research opportunities of this novel creative practice in the field of Human-Computer Interaction (HCI). The paper concludes with a discussion of broader implications of prompt engineering from the perspective of Human-AI Interaction (HAI) in future applications beyond the use case of text-to-image generation and AI generated art.
15 pages
References in corpus (11)
- On the Opportunities and Risks of Foundation Models
- Evaluating Large Language Models Trained on Code
- Zero-Shot Text-to-Image Generation
- Imagen Video: High Definition Video Generation with Diffusion Models
- Make-A-Video: Text-to-Video Generation without Text-Video Data
- The Creativity of Text-to-Image Generation
- Augmented Language Models: a Survey
- CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
- Text-Guided Synthesis of Artistic Images with Retrieval-Augmented Diffusion Models
- DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models
- Is Writing Prompts Really Making Art?
Cited by in corpus (11)
- Using Text-to-Image Generation for Architectural Design Ideation
- Mapping the Ethics of Generative AI: A Comprehensive Scoping Review
- PromptMagician: Interactive Prompt Engineering for Text-to-Image Creation
- Perceptions and Realities of Text-to-Image Generation
- Prompt Evolution for Generative AI: A Classifier-Guided Approach
- PromptMap: An Alternative Interaction Style for AI-Based Image Generation
- A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
- The Cow of Rembrandt - Analyzing Artistic Prompt Interpretation in Text-to-Image Models
- Stable diffusion models reveal a persisting human and AI gap in visual creativity
- Varif.ai to Vary and Verify User-Driven Diversity in Scalable Image Generation
- SSP: A Simple and Safe automatic Prompt engineering method towards realistic image synthesis on LVM