This is the first controlled test of whether saliency-map instability adds failure-detection information beyond predictive uncertainty in mammography.
1. Across five CNN architectures and three settings (cropped-lesion CBIS-DDSM, multi-seed full-image CBIS-DDSM, and an external CMMD cohort), predictive entropy reliably detects classification errors (error-detection AUC 0.63–0.70) and supports effective selective prediction.
2. Using nested likelihood-ratio tests with paired bootstrapping and Benjamini–Hochberg correction, we show that saliency-map instability performs near chance on CBIS-DDSM (AUC approximately 0.50) and adds no significant information beyond predictive uncertainty in any full-image model, replicated across three seeds.
3. A controlled cropped-versus-full ablation shows the modest CMMD gains (AUC +0.015 to 0.054) are not reproduced on full-image CBIS-DDSM, pointing to dataset- or domain-specific factors rather than image format.
