paper

The effects of data preprocessing on probability of default model fairness

arXiv:2408.15452 · doi:10.30574/wjaets.2024.12.2.0354

Abstract

In the context of financial credit risk evaluation, the fairness of machine learning models has become a critical concern, especially given the potential for biased predictions that disproportionately affect certain demographic groups. This study investigates the impact of data preprocessing, with a specific focus on Truncated Singular Value Decomposition (SVD), on the fairness and performance of probability of default models. Using a comprehensive dataset sourced from Kaggle, various preprocessing techniques, including SVD, were applied to assess their effect on model accuracy, discriminatory power, and fairness.

The effects of data preprocessing on probability of default model fairness · wovepaper