An Explainable Meta-Ensemble Framework for Leakage-Free Active Tuberculosis Diagnosis in High-Dimensional Transcriptomics

Tuberculosis (TB) remains a leading infectious cause of death, and host blood transcriptomics offers a non-invasive route to detecting active disease. Applying machine learning here is hard because the genes measured vastly outnumber the patients (n≪p), inviting overfitting, and because many pipelines select biomarkers on the full dataset before splitting, leaking test information and inflating mean accuracy. We present a zero-leakage, explainable meta-ensemble for active-TB diagnosis in which every data-dependent step, standardization, ANOVA biomarker purification, and SMOTE balancing, is confined inside the training folds of a stratified 10-fold cross-validation, so the reported metrics reflect genuinely unseen patients. The classifier is a soft-voting ensemble pairing LightGBM, a sequential gradient booster, with an Extra-Trees classifier that shields against high-dimensional noise. Under this protocol, it reaches a stable 95.03% mean accuracy (ROC-AUC 0.9884). Game-theoretic SHAP analysis shows that the decision is driven by immunological markers, the HLA and MHC genes, rather than noise. The top 76 SHAP ranked genes were then analysed through pathway and Gene Ontology enrichment, protein–protein interaction (PPI) and hub gene analysis, transcription-factor and microRNA inference, and drug–protein mining, anchoring the panel to antigen processing and presentation, with HLA-A/B/C and TAPBP as central hubs. The study shows how a disciplined, interpretable pipeline turns high-dimensional TB transcriptomics into trustworthy predictions and biologically grounded biomarkers.