Fairness in Credit Scoring: Assessment, Implementation and Profit Implications
arXiv:2103.01907 · doi:10.1016/j.ejor.2021.06.023
Abstract
The rise of algorithmic decision-making has spawned much research on fair machine learning (ML). Financial institutions use ML for building risk scorecards that support a range of credit-related decisions. Yet, the literature on fair ML in credit scoring is scarce. The paper makes three contributions. First, we revisit statistical fairness criteria and examine their adequacy for credit scoring. Second, we catalog algorithmic options for incorporating fairness goals in the ML model development pipeline. Last, we empirically compare different fairness processors in a profit-oriented credit scoring context using real-world data. The empirical results substantiate the evaluation of fairness measures, identify suitable options to implement fair credit scoring, and clarify the profit-fairness trade-off in lending decisions. We find that multiple fairness criteria can be approximately satisfied at once and recommend separation as a proper criterion for measuring the fairness of a scorecard. We also find fair in-processors to deliver a good balance between profit and fairness and show that algorithmic discrimination can be reduced to a reasonable level at a relatively low cost. The codes corresponding to the paper are available on GitHub.
Accepted to European Journal of Operational Research
References in corpus (3)
Cited by in corpus (11)
- A Fused Large Language Model for Predicting Startup Success
- Whither Bias Goes, I Will Go: An Integrative, Systematic Review of Algorithmic Bias Mitigation
- Algorithmic decision making methods for fair credit scoring
- Bridging the gap: Towards an Expanded Toolkit for AI-driven Decision-Making in the Public Sector
- Equalizing Credit Opportunity in Algorithms: Aligning Algorithmic Fairness Research with U.S. Fair Lending Regulation
- Enforcing Group Fairness in Algorithmic Decision Making: Utility Maximization Under Sufficiency
- Resolving Ethics Trade-offs in Implementing Responsible AI
- NLPGuard: A Framework for Mitigating the Use of Protected Attributes by NLP Classifiers
- Supervised Feature Compression based on Counterfactual Analysis
- Counterfactual Situation Testing: From Single to Multidimensional Discrimination
- FairGridSearch: A Framework to Compare Fairness-Enhancing Models