Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
arXiv:2209.08891 · doi:10.1613/jair.1.15388
Abstract
Models for text-to-image synthesis, such as DALL-E~2 and Stable Diffusion, have recently drawn a lot of interest from academia and the general public. These models are capable of producing high-quality images that depict a variety of concepts and styles when conditioned on textual descriptions. However, these models adopt cultural characteristics associated with specific Unicode scripts from their vast amount of training data, which may not be immediately apparent. We show that by simply inserting single non-Latin characters in a textual description, common models reflect cultural stereotypes and biases in their generated images. We analyze this behavior both qualitatively and quantitatively, and identify a model's text encoder as the root cause of the phenomenon. Additionally, malicious users or service providers may try to intentionally bias the image generation to create racist stereotypes by replacing Latin characters with similarly-looking characters from non-Latin scripts, so-called homoglyphs. To mitigate such unnoticed script attacks, we propose a novel homoglyph unlearning method to fine-tune a text encoder, making it robust against homoglyph manipulations.
Published in the Journal of Artificial Intelligence Research (JAIR)
References in corpus (21)
- Explaining and Harnessing Adversarial Examples
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Zero-Shot Text-to-Image Generation
- LAION-5B: An open large-scale dataset for training next generation image-text models
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Do ImageNet Classifiers Generalize to ImageNet?
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
- Easily Accessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale
- Multimodal datasets: misogyny, pornography, and malignant stereotypes
- A very preliminary analysis of DALL-E 2
- Tackling Algorithmic Disability Discrimination in the Hiring Process: An Ethical, Legal and Technical Analysis
- Stable Bias: Analyzing Societal Representations in Diffusion Models
- Testing Relational Understanding in Text-Guided Image Generation
- Fair Diffusion: Instructing Text-to-Image Generation Models on Fairness
- Adversarial Attacks on Image Generation With Made-Up Words
- Improving Sample Quality of Diffusion Models Using Self-Attention Guidance
- LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing
- AltDiffusion: A Multilingual Text-to-Image Diffusion Model