Bias and Unfairness in Information Retrieval Systems: New Challenges in the LLM Era
arXiv:2404.11457 · doi:10.1145/3637528.3671458
Abstract
With the rapid advancements of large language models (LLMs), information retrieval (IR) systems, such as search engines and recommender systems, have undergone a significant paradigm shift. This evolution, while heralding new opportunities, introduces emerging challenges, particularly in terms of biases and unfairness, which may threaten the information ecosystem. In this paper, we present a comprehensive survey of existing works on emerging and pressing bias and unfairness issues in IR systems when the integration of LLMs. We first unify bias and unfairness issues as distribution mismatch problems, providing a groundwork for categorizing various mitigation strategies through distribution alignment. Subsequently, we systematically delve into the specific bias and unfairness issues arising from three critical stages of LLMs integration into IR systems: data collection, model development, and result evaluation. In doing so, we meticulously review and analyze recent literature, focusing on the definitions, characteristics, and corresponding mitigation strategies associated with these issues. Finally, we identify and highlight some open problems and challenges for future work, aiming to inspire researchers and stakeholders in the IR field and beyond to better understand and mitigate bias and unfairness issues of IR in this LLM era. We also consistently maintain a GitHub repository for the relevant papers and resources in this rising direction at https://github.com/KID-22/LLM-IR-Bias-Fairness-Survey.
KDD 2024 Tutorial&Survey; Tutorial Website: https://llm-ir-bias-fairness.github.io/
References in corpus (3)
Cited by in corpus (12)
- Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges
- SPRec: Self-Play to Debias LLM-based Recommendation
- Rankers, Judges, and Assistants: Towards Understanding the Interplay of LLMs in Information Retrieval Evaluation
- IdeaSynth: Iterative Research Idea Development Through Evolving and Composing Idea Facets with Literature-Grounded Feedback
- Capturing research literature attitude towards Sustainable Development Goals: an LLM-based topic modeling approach
- Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-In-The-Loop LLM
- Evaluating Large Language Models in Code Generation: INFINITE Methodology for Defining the Inference Index
- NExT-Search: Rebuilding User Feedback Ecosystem for Generative AI Search
- The Effects of Demographic Instructions on LLM Personas
- Biases in LLM-Generated Musical Taste Profiles for Recommendation
- Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
- How Can AI Augment Access to Justice? Public Defenders' Perspectives on AI Adoption