View-Aware Reliability Framework for Pneumonia Diagnosis with Calibration Portability, Selective Risk, and XAI Auditing

Deep learning models for pneumonia detection from chest radiographs are usually assessed with discrimination metrics, yet comparable discrimination does not guarantee that predicted probabilities remain trustworthy across acquisition views or institutions. This work evaluates a DenseNet121 pneumonia classifier under anteroposterior (AP) and posteroanterior (PA) view variation and cross-dataset shift. A patient-level cohort from NIH ChestX-ray14 was used for development and internal testing, and an independent CheXpert cohort was used for external testing with no retraining, calibration refitting, or threshold adjustment. Global temperature scaling was compared with AP/PA-specific temperature scaling, predictive entropy supported uncertainty-guided selective prediction, and Grad-CAM provided a qualitative reliability audit. The classifier reached AUROCs of 0.7958 on NIH and 0.7859 on CheXpert, indicating relatively stable ranking behavior. Reliability was more sensitive to the shift. View-aware calibration lowered internal expected calibration error from 0.0460 to 0.0361, but the source-fitted temperatures did not improve on raw CheXpert probabilities (0.0822 versus 0.0813). The entropy threshold preserved coverage externally (90.10% to 88.67%) while the area under the riskcoverage curve rose from 0.1285 to 0.1564. PA radiographs showed lower discrimination and poorer calibration than AP radiographs in both cohorts. These results show that calibration and selective reliability can degrade while AUROC appears stable, supporting view- and domain-aware reliability evaluation before cross-dataset deployment.