BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
arXiv:2101.11718 · doi:10.1145/3442188.3445924
Abstract
Recent advances in deep learning techniques have enabled machines to generate cohesive open-ended text when prompted with a sequence of words as context. While these models now empower many downstream applications from conversation bots to automatic storytelling, they have been shown to generate texts that exhibit social biases. To systematically study and benchmark social biases in open-ended language generation, we introduce the Bias in Open-Ended Language Generation Dataset (BOLD), a large-scale dataset that consists of 23,679 English text generation prompts for bias benchmarking across five domains: profession, gender, race, religion, and political ideology. We also propose new automated metrics for toxicity, psycholinguistic norms, and text gender polarity to measure social biases in open-ended text generation from multiple angles. An examination of text generated from three popular language models reveals that the majority of these models exhibit a larger social bias than human-written Wikipedia text across all domains. With these results we highlight the need to benchmark biases in open-ended language generation and caution users of language generation models on downstream tasks to be cognizant of these embedded prejudices.
Cited by in corpus (25)
- Black-Box Access is Insufficient for Rigorous AI Audits
- LaMPost: Design and Evaluation of an AI-assisted Email Writing Prototype for Adults with Dyslexia
- "I'm fully who I am": Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation
- A Security Risk Taxonomy for Prompt-Based Interaction With Large Language Models
- A Scoping Study of Evaluation Practices for Responsible AI Tools: Steps Towards Effectiveness Evaluations
- Clinical Insights: A Comprehensive Review of Language Models in Medicine
- Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
- KnowledgeVIS: Interpreting Language Models by Comparing Fill-in-the-Blank Prompts
- LLM-Driven Robots Risk Enacting Discrimination, Violence, and Unlawful Actions
- Lazy Data Practices Harm Fairness Research
- On LLM Wizards: Identifying Large Language Models' Behaviors for Wizard of Oz Experiments
- Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
- Religious Bias Landscape in Language and Text-to-Image Models: Analysis, Detection, and Debiasing Strategies
- Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge
- Laissez-Faire Harms: Algorithmic Biases in Generative Language Models
- LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases
- A Matter of Perspective(s): Contrasting Human and LLM Argumentation in Subjective Decision-Making on Subtle Sexism
- Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language Models
- COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
- Are Large Language Models Really Bias-Free? Jailbreak Prompts for Assessing Adversarial Robustness to Bias Elicitation
- LLM Bias Evaluation: Gender, Racial, and Age Disparities in Occupational and Crime Scenarios
- Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement
- No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
- CFL: Causally Fair Language Models Through Token-level Attribute Controlled Generation
- Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study