Two Public CKD Benchmarks Are Not Independent: A Reproducible Audit of Leakage, Cohort Separability, and Pseudo-External Validatio

This work provides record-level evidence that two widely-used public CKD benchmark datasets (the UCI 2015 Indian release and the 2020 Bangladeshi release) are not independent: exact-match bipartite record linkage shows all 200 patients in the 2015 dataset also appear in the 2020 dataset, contradicting their documented separate provenance. The paper demonstrates that models trained on one and “externally validated” on the other are not measuring transportability but rather leakage — near-ceiling ROC-AUC performance (up to 1.000) is shown to result from post-diagnosis outcome-correlated variables rather than genuine predictive power, dropping to a more realistic 0.909 once such variables are removed. The study contributes a reusable, dataset-agnostic auditing methodology (a provenance gate, feature taxonomy, and leakage-prevention framework) for detecting duplicate contamination and incorporation bias in clinical prediction benchmarks generally.