Calibrated and Explainable Credit Risk Assessment: Integrating SHAP Analysis with Feature Stability Validation

Machine learning has substantially improved the predictive performance of credit risk assessment, yet high pre- dictive accuracy alone is insufficient for deployment in regu- lated lending, where reliable probability estimates, interpretable decisions, and consistent explanations are equally important. Existing studies typically evaluate these properties independently, leaving uncertainty about whether a model can satisfy all three simultaneously. This paper proposes an integrated evaluation framework that jointly assesses discrimination, probability cali- bration, and explanation stability with statistical confidence. Us- ing the German Credit dataset, we compare Logistic Regression, Random Forest, and LightGBM. Model performance is evaluated using ROC-AUC, the Brier score, Expected Calibration Error (ECE), bootstrap confidence intervals, and paired significance tests, while TreeSHAP explanations are validated across five cross-validation folds using feature-wise coefficients of variation. Experimental results show that LightGBM achieves the best overall discrimination and calibration while maintaining stable explanations for its most influential risk factors. However, its performance advantage over Random Forest is not statistically significant. We further demonstrate that synthetic oversampling improves recall but consistently degrades probability calibration, whereas cost-sensitive threshold optimization achieves compara- ble gains without distorting predicted probabilities. In addition, a fairness audit reveals demographic disparities that are not reflected in feature attribution alone. These findings highlight that discrimination, calibration, explanation stability, and fairness should be evaluated together, providing a more reliable basis for selecting and validating credit risk models for practical deployment. Index Ter