Leakage-Free, Calibrated, and Explainable Machine Learning for Heart-Disease Screening from Large-Scale BRFSS Survey Data

Our contributions are as follows.
• A leakage-free benchmark. Under honest resampling,
the task is genuinely hard (test ROC–AUC ≈ 0.81), and
stacking does not beat a well-regularized linear baseline.
• A screening-oriented evaluation. We report ROC–AUC
with sensitivity, specificity, and MCC, add isotonic cal-
ibration, and make the sensitivity/specificity trade-off
explicit through threshold analysis.
• Trust and transparency. We audit fairness across sex,
race, and age, show that age-stratified thresholds reduce
the age-related inequity, quantify uncertainty, and ex-
plain the model with permutation importance, SHAP, and
LIME.