An Explainable AI Approach for Bus Safety Risk Classification in Data-Scarce Transport Networks

Bus overtaking and reckless driving are important
road-safety concerns in data-scarce low- and middle-income
cities. This paper develops a machine-learning and explainable-AI
(XAI) framework for Dhaka, Bangladesh, where structured bus
telematics, digital driver logs, mechanical-condition histories, and
bus-specific incident labels are not readily available. The central
contribution is the Dhaka Bus Driving Risk Dataset (DBDRD), a
domain-grounded synthetic benchmark containing 12,000 records
and 43 variables across six feature groups. Its generation
process combines context-specific distributions, explicit conditional
dependencies, a 12-component literature-calibrated risk score,
Gaussian perturbation, percentile-based ordinal labeling, and
controlled missing-at-random noise. Random Forest, XGBoost,
LightGBM, and radial-basis-function SVM were evaluated using
a leakage-controlled 70/15/15 split and cross-validated tuning.
XGBoost achieved the strongest overall test performance with
76.4% accuracy, 0.761 macro F1, 0.767 weighted F1, and 0.937
one-versus-rest ROC-AUC. SHAP analysis identified operational
risk, aggression, compliance risk, and speed excess as the dominant
predictors and revealed compound behavioral-mechanical effects
in Critical-risk cases. The results establish a reproducible synthetic
benchmark and interpretable modeling pipeline, while external
validity remains contingent on future validation with real Dhaka
bus data.