Toward Trustworthy Stroke Risk Prediction: A Balanced, Explainable and Externally Validated Machine Learning Framework

This study presents an interpretable and trustworthy stroke risk prediction framework that integrates Boruta–mRMR feature selection, training-only Borderline-SMOTE, comparative evaluation of nine classifiers, and an optimized CatBoost model with an OOF F2-based threshold. The framework further combines SHAP, LIME, and counterfactual explanations with fairness, calibration, robustness, bootstrap-based reliability, and cross-dataset external validation.