Beyond Point Estimates: Stability-Aware Interpretable Machine Learning for Drug-Induced Autoimmunity Prediction

This paper addresses both gaps directly on a public UCI
dataset, the RDKit-descriptor release underlying InterDIA.
Our contributions are:
1) A leakage-safe, nested cross-validation benchmark of
six algorithm families (linear, bagged-tree, two boosting
variants, and an imbalance-aware ensemble), with
algorithm ranking corroborated by a Friedman/Nemenyi
statistical test, and confirmed on the official heldout
test set.
2) A bootstrapped feature-selection stability analysis using
Jaccard similarity and Nogueira’s chance-corrected
stability index, quantifying how much ”important”
descriptors shift under resampling in this highdimensional,
small-sample regime.
3) A systematic SHAP-vs-LIME cross-validation, comparing
global and per-instance rank agreement and top-k
feature overlap, to test whether aggregate XAI agreement
masks local disagreement.
4) A descriptor-space matched-pair mechanistic sanity
check, inspired by InterDIA’s structurally matched case
studies and by matched molecular pair analysis in
medicinal chemistry, used here as a proxy robustness
probe rather than a validated structural-analogue
analysis.