Large-scale Text-to-Image Generation Models for Visual Artists' Creative Works
arXiv:2210.08477 · doi:10.1145/3581641.3584078
Abstract
Large-scale Text-to-image Generation Models (LTGMs) (e.g., DALL-E), self-supervised deep learning models trained on a huge dataset, have demonstrated the capacity for generating high-quality open-domain images from multi-modal input. Although they can even produce anthropomorphized versions of objects and animals, combine irrelevant concepts in reasonable ways, and give variation to any user-provided images, we witnessed such rapid technological advancement left many visual artists disoriented in leveraging LTGMs more actively in their creative works. Our goal in this work is to understand how visual artists would adopt LTGMs to support their creative works. To this end, we conducted an interview study as well as a systematic literature review of 72 system/application papers for a thorough examination. A total of 28 visual artists covering 35 distinct visual art domains acknowledged LTGMs' versatile roles with high usability to support creative works in automating the creation process (i.e., automation), expanding their ideas (i.e., exploration), and facilitating or arbitrating in communication (i.e., mediation). We conclude by providing four design guidelines that future researchers can refer to in making intelligent user interfaces using LTGMs.
15 pages, 3 figures
References in corpus (5)
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
- The Impact of Multiple Parallel Phrase Suggestions on Email Input and Composition Behaviour of Native and Non-Native English Writers
- Perfection Not Required? Human-AI Partnerships in Code Translation
- Better Together? An Evaluation of AI-Supported Code Translation
- SGToolkit: An Interactive Gesture Authoring Toolkit for Embodied Conversational Agents
Cited by in corpus (22)
- PromptCharm: Text-to-Image Generation through Multi-modal Prompting and Refinement
- AIdeation: Designing a Human-AI Collaborative Ideation System for Concept Designers
- Sketchar: Supporting Character Design and Illustration Prototyping Using Generative AI
- PlantoGraphy: Incorporating Iterative Design Process into Generative Artificial Intelligence for Landscape Rendering
- From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape
- Towards A Diffractive Analysis of Prompt-Based Generative AI
- From Paper to Card: Transforming Design Implications with Generative AI
- AI-Assisted Causal Pathway Diagram for Human-Centered Design
- AI Rivalry as a Craft: How Resisting and Embracing Generative AI Reshape Writing Professions
- Situating the social issues of image generation models in the model life cycle: a sociotechnical approach
- "Salt is the Soul of Hakka Baked Chicken": Reimagining Traditional Chinese Culinary ICH for Modern Contexts Without Losing Tradition
- Expandora: Broadening Design Exploration with Text-to-Image Model
- MemoVis: A GenAI-Powered Tool for Creating Companion Reference Images for 3D Design Feedback
- PromptMap: An Alternative Interaction Style for AI-Based Image Generation
- What Social Media Use Do People Regret? An Analysis of 34K Smartphone Screenshots with Multimodal LLM
- A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
- Do It For Me vs. Do It With Me: Investigating User Perceptions of Different Paradigms of Automation in Copilots for Feature-Rich Software
- PosterMate: Audience-driven Collaborative Persona Agents for Poster Design
- GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment Design
- Understanding Collaboration between Professional Designers and Decision-making AI: A Case Study in the Workplace
- Offline Evaluation of Set-Based Text-to-Image Generation
- Empowering Children to Create AI-Enabled Augmented Reality Experiences