Ensemble Deep Learning Framework for Calibrated and Explainable Diabetic Retinopathy Grading on RetinaMNIST

The importance of this study is that it aims to shift the focus from achieving grading accuracy in DR towards a more reliable, clinically meaningful and interpretable grading system. The proposed four-model ensemble obtained a QWK score of 0.813 and macro-AUC of 0.908, which are better than the previously reported baseline results on RetinaMNIST and achieved better ordinal and discriminative results.

Most importantly, the study merges the ordinal-aware evaluation with the explainability, calibration, uncertainty analysis and structured error analysis of Grad-CAM. This not only provides an assessment of the accuracy of the model, but also the validity of the level of confidence, whether the model targets clinically relevant retinal lesions, and where potentially harmful errors (e.g. under-grading PDR) may arise.

Thus, the key contribution is a more holistic perspective on the assessment of AI-based DR grading, focusing not just on accuracy, but also on clinical trust and risk awareness.