A Leakage-Aware Ordinal Machine Learning and Explainable AI Framework for Hidden Replication, Label Ambiguity, and Performance Inflation in Maternal Health Risk Prediction

This paper proposes four significant contributions:
1) We quantify the scale of duplication and label ambiguity and its practical consequence for standard train/test splitting.
2) We directly measure the resulting performance inflation by comparing conventional cross-validation against exact-row-grouped and physiological-profile-grouped cross-validation for two widely used models.
3) We benchmark a broad family of ordinal-aware and conventional classifiers under leakage-aware, repeated nested cross-validation, including a two-threshold ordinal decomposition and a CORAL-style ordinal neural network, with full calibration and conformal-prediction analysis.
4) We evaluate the stability of SHAP-based explanations across random seeds, treating explanation reliability itself as an empirical outcome rather than an illustrative afterthought.