FairCoder: Probing LLM Bias in High-Stakes Decision Making via Coding Tasks
arXiv:2501.05396
The paper introduces FairCoder, a benchmark that uses coding tasks to detect implicit social bias in large language models when they are applied to high‑stakes decisions like hiring, college admissions, and healthcare, and proposes a new metric, FairScore, to evaluate both refusal behavior and outcome diversity.
Abstract
Large language models (LLMs) are increasingly used in high-stakes decisions such as hiring and college admissions, making their social bias a critical concern. While LLMs are trained to refuse explicitly biased requests, bias can be leaked implicitly during LLM planning and reasoning process. As code becomes the primary medium for LLM internal logic-writing, we introduce FairCoder, a benchmark that frames decision-making as coding tasks to systematically probe LLM bias across employment, education, and healthcare domains, covering multiple fairness definitions. Considering that existing metrics may fail when LLMs frequently refuse the request, we propose FairScore, a metric that jointly captures refusal behavior and group-level outcome diversity. Experiments with a 1k-sample dataset on powerful LLMs reveal consistent and previously underexplored bias patterns, such as prioritizing applicants from high-income families in college admissions. Our findings highlight the risks of deploying LLMs as decision-making agents and provide a comprehensive evaluation framework for future research.