Is ChatGPT Fair for Recommendation? Evaluating Fairness in Large Language Model Recommendation
arXiv:2305.07609 · doi:10.1145/3604915.3608860
Abstract
The remarkable achievements of Large Language Models (LLMs) have led to the emergence of a novel recommendation paradigm -- Recommendation via LLM (RecLLM). Nevertheless, it is important to note that LLMs may contain social prejudices, and therefore, the fairness of recommendations made by RecLLM requires further investigation. To avoid the potential risks of RecLLM, it is imperative to evaluate the fairness of RecLLM with respect to various sensitive attributes on the user side. Due to the differences between the RecLLM paradigm and the traditional recommendation paradigm, it is problematic to directly use the fairness benchmark of traditional recommendation. To address the dilemma, we propose a novel benchmark called Fairness of Recommendation via LLM (FaiRLLM). This benchmark comprises carefully crafted metrics and a dataset that accounts for eight sensitive attributes1 in two recommendation scenarios: music and movies. By utilizing our FaiRLLM benchmark, we conducted an evaluation of ChatGPT and discovered that it still exhibits unfairness to some sensitive attributes when generating recommendations. Our code and dataset can be found at https://github.com/jizhi-zhang/FaiRLLM.
Accepted by Recsys 2023 (Short). Typo corrections
References in corpus (7)
- A Survey of Large Language Models
- TALLRec: An Effective and Efficient Tuning Framework to Align Large Language Model with Recommendation
- User-oriented Fairness in Recommendation
- Personalized Counterfactual Fairness in Recommendation
- Toward Pareto Efficient Fairness-Utility Trade-off inRecommendation through Reinforcement Learning
- An Audit of Misinformation Filter Bubbles on YouTube: Bubble Bursting and Recent Behavior Changes
- Generative Recommendation: Towards Next-generation Recommender Paradigm
Cited by in corpus (9)
- TALLRec: An Effective and Efficient Tuning Framework to Align Large Language Model with Recommendation
- Recommender Systems in the Era of Large Language Models (LLMs)
- Bias and Unfairness in Information Retrieval Systems: New Challenges in the LLM Era
- SPRec: Self-Play to Debias LLM-based Recommendation
- Adaptive In-Context Learning with Large Language Models for Bundle Generation
- LangFair: A Python Package for Assessing Bias and Fairness in Large Language Model Use Cases
- Fairness Definitions in Language Models Explained
- Quantitative Fairness -- A Framework For The Design Of Equitable Cybernetic Societies
- Stairway to Fairness: Connecting Group and Individual Fairness