GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models
arXiv:2407.21001 · doi:10.3233/FAIA240555
Abstract
Vision-language models (VLMs) are intensively used in many downstream tasks, including those requiring assessments of individuals appearing in the images. While VLMs perform well in simple single-person scenarios, in real-world applications, we often face complex situations in which there are persons of different genders doing different activities. We show that in such cases, VLMs are biased towards identifying the individual with the expected gender (according to ingrained gender stereotypes in the model or other forms of sample selection bias) as the performer of the activity. We refer to this bias in associating an activity with the gender of its actual performer in an image or text as the Gender-Activity Binding (GAB) bias and analyze how this bias is internalized in VLMs. To assess this bias, we have introduced the GAB dataset with approximately 5500 AI-generated images that represent a variety of activities, addressing the scarcity of real-world images for some scenarios. To have extensive quality control, the generated images are evaluated for their diversity, quality, and realism. We have tested 12 renowned pre-trained VLMs on this dataset in the context of text-to-image and image-to-text retrieval to measure the effect of this bias on their predictions. Additionally, we have carried out supplementary experiments to quantify the bias in VLMs' text encoders and to evaluate VLMs' capability to recognize activities. Our experiments indicate that VLMs experience an average performance decline of about 13.2% when confronted with gender-activity binding bias.
References in corpus (18)
- Denoising Diffusion Probabilistic Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- CoCa: Contrastive Captioners are Image-Text Foundation Models
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- FILIP: Fine-grained Interactive Language-Image Pre-Training
- EVA-CLIP: Improved Training Techniques for CLIP at Scale
- When and why vision-language models behave like bags-of-words, and what to do about it?
- Evaluating CLIP: Towards Characterization of Broader Capabilities and Downstream Implications
- Multimodal Foundation Models: From Specialists to General-Purpose Assistants
- Debiasing Vision-Language Models via Biased Prompts
- VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations
- Vision Language Models in Autonomous Driving: A Survey and Outlook
- Survey of Social Bias in Vision-Language Models
- VisoGender: A dataset for benchmarking gender bias in image-text pronoun resolution