Fairness Definitions in Language Models Explained
arXiv:2407.18454 · doi:10.1002/widm.70063
Abstract
Language Models (LMs) have demonstrated exceptional performance across various Natural Language Processing (NLP) tasks. Despite these advancements, LMs can inherit and amplify societal biases related to sensitive attributes such as gender and race, limiting their adoption in real-world applications. Therefore, fairness has been extensively explored in LMs, leading to the proposal of various fairness notions. However, the lack of clear agreement on which fairness definition to apply in specific contexts and the complexity of understanding the distinctions between these definitions can create confusion and impede further progress. To this end, this paper proposes a systematic survey that clarifies the definitions of fairness as they apply to LMs. Specifically, we begin with a brief introduction to LMs and fairness in LMs, followed by a comprehensive, up-to-date overview of existing fairness notions in LMs and the introduction of a novel taxonomy that categorizes these concepts based on their transformer architecture: encoder-only, decoder-only, and encoder-decoder LMs. We further illustrate each definition through experiments, showcasing their practical implications and outcomes. Finally, we discuss current research challenges and open questions, aiming to foster innovative ideas and advance the field. The repository is publicly available online at https://github.com/vanbanTruong/Fairness-in-Large-Language-Models/tree/main/definitions.
References in corpus (9)
- Semantics derived automatically from language corpora contain human-like biases
- Fairness in Machine Learning: A Survey
- Gender bias and stereotypes in Large Language Models
- Fairness in Machine Learning
- Bias in Bios: A Case Study of Semantic Representation Bias in a High-Stakes Setting
- Masked Language Model Scoring
- Is ChatGPT Fair for Recommendation? Evaluating Fairness in Large Language Model Recommendation
- Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases
- Predictive Biases in Natural Language Processing Models: A Conceptual Framework and Overview