Easily Accessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale
arXiv:2211.03759 · doi:10.1145/3593013.3594095
Abstract
Machine learning models that convert user-written text descriptions into images are now widely available online and used by millions of users to generate millions of images a day. We investigate the potential for these models to amplify dangerous and complex stereotypes. We find a broad range of ordinary prompts produce stereotypes, including prompts simply mentioning traits, descriptors, occupations, or objects. For example, we find cases of prompting for basic traits or social roles resulting in images reinforcing whiteness as ideal, prompting for occupations resulting in amplification of racial and gender disparities, and prompting for objects resulting in reification of American norms. Stereotypes are present regardless of whether prompts explicitly mention identity and demographic language or avoid such language. Moreover, stereotypes persist despite mitigation strategies; neither user attempts to counter stereotypes by requesting images with specific counter-stereotypes nor institutional attempts to add system ``guardrails'' have prevented the perpetuation of stereotypes. Our analysis justifies concerns regarding the impacts of today's models, presenting striking exemplars, and connecting these findings with deep insights into harms drawn from social scientific and humanist disciplines. This work contributes to the effort to shed light on the uniquely complex biases in language-vision models and demonstrates the ways that the mass deployment of text-to-image generation models results in mass dissemination of stereotypes and resulting harms.
FAccT 2023 paper. The published version is available at 10.1145/3593013.3594095
References in corpus (3)
Cited by in corpus (20)
- Using Text-to-Image Generation for Architectural Design Ideation
- AI's Regimes of Representation: A Community-centered Study of Text-to-Image Models in South Asia
- Not My Voice! A Taxonomy of Ethical and Safety Harms of Speech Generators
- Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
- 'Person' == Light-skinned, Western Man, and Sexualization of Women of Color: Stereotypes in Stable Diffusion
- From tools to thieves: Measuring and understanding public perceptions of AI through crowdsourced metaphors
- Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
- HarmonyCut: Supporting Creative Chinese Paper-cutting Design with Form and Connotation Harmony
- Un-Straightening Generative AI: How Queer Artists Surface and Challenge the Normativity of Generative AI Models
- Situating the social issues of image generation models in the model life cycle: a sociotechnical approach
- Who's in and who's out? A case study of multimodal CLIP-filtering in DataComp
- Auditing Gender Presentation Differences in Text-to-Image Models
- Religious Bias Landscape in Language and Text-to-Image Models: Analysis, Detection, and Debiasing Strategies
- Vox Populi, Vox AI? Using Language Models to Estimate German Public Opinion
- Learning AI Auditing: A Case Study of Teenagers Auditing a Generative AI Model
- Ontologies in Design: How Imagining a Tree Reveals Possibilites and Assumptions in Large Language Models
- Social Perception of Faces in a Vision-Language Model
- Understanding Gender Bias in AI-Generated Product Descriptions
- Interactive Discovery and Exploration of Visual Bias in Generative Text-to-Image Models
- Varif.ai to Vary and Verify User-Driven Diversity in Scalable Image Generation