The Infinite Index: Information Retrieval on Generative Text-To-Image Models
arXiv:2212.07476 · doi:10.1145/3576840.3578327
Abstract
Conditional generative models such as DALL-E and Stable Diffusion generate images based on a user-defined text, the prompt. Finding and refining prompts that produce a desired image has become the art of prompt engineering. Generative models do not provide a built-in retrieval model for a user's information need expressed through prompts. In light of an extensive literature review, we reframe prompt engineering for generative models as interactive text-based retrieval on a novel kind of "infinite index". We apply these insights for the first time in a case study on image generation for game design with an expert. Finally, we envision how active learning may help to guide the retrieval of generated images.
Final version for CHIIR 2023
References in corpus (6)
- Diffusion Models Beat GANs on Image Synthesis
- DreamFusion: Text-to-3D using 2D Diffusion
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
- Imagen Video: High Definition Video Generation with Diffusion Models
- Autoregressive Search Engines: Generating Substrings as Document Identifiers
- Lecture Notes on Neural Information Retrieval