An AI-Enhanced Four-Class Stacking Ensemble Framework for Explainable Income Tax Fraud Detection Using Calibrated NBR IT-10B Structure

Revenue collection is severely hampered by income
tax fraud, especially in developing nations like Bangladesh. The
National Board of Revenue (NBR) has to deal with issues including severely skewed fraud statistics and a dearth of instruments
that can recognize various forms of fraud from the nation’s own
tax return system. The majority of current machine learning
research is mostly on binary fraud detection and offers scant
justification for their forecasts. Based on Bangladesh’s IT-10B
individual tax return structure and the 2025–2026 tax slabs, this
paper suggests a four-class income tax fraud classification scheme
that covers No Fraud, Underreporting, Inflated Deductions, and
False Credits. With the assistance of a practicing tax attorney,
a dataset of 2,498 records was created from anonymized tax
files and verified. Experts in the field were consulted in order to
further verify the fraud typology guidelines. A weighted Borda
aggregation of RFECV, LightGBM feature significance, Pearson
correlation, and Mutual Information was used to decrease the
initial 62 features produced by the proposed framework to
25. To solve class imbalance, SVMSMOTE was solely used on
training data. A pruned stacking ensemble was used to merge
five models (SVM, Random Forest, ANN, CatBoost, and Logistic
Regression), with Logistic Regression serving as the meta-learner.
The suggested model obtained a macro F1-score of 0.9502 and
an accuracy of 96.53% on the locked test set.It fared better
than the soft-voting ensemble (0.9272), the best individual model,
SVM (macro F1: 0.9343), and four re-implemented literature
baselines tested on the same dataset and test split. The model’s
general behavior and class-specific predictions were explained
using SHAP and LIME.