Road accidents have no single factor to take place.
There are many different variables that affect road accidents
including weather, road, traffic, enforcement, drivers, and many
others. The interaction between these variables influences the
number of road accidents. In addition, there are few crash
database available for Bangladesh, and that’s why we created a
synthetic dataset to conduct our analysis. Our data set includes
9,000 records with 28 different columns, each of which represents
the scenario of the district under consideration in the given
month with particular road, weather, and traffic conditions. Five
columns were dropped since they did not exist before the crash
event happened, giving 22 predictors in total. Random Forest,
Extra Trees, Gradient Boosting, and soft voting were performed
in addition to the previous techniques. The base technique was
Logistic Regression. The maximum accuracy score of 83.00% was
attained by Gradient Boosting. Soft voting gave 82.83%. Linear
baseline had 74.44% accuracy score. The macro-average precision,
recall, and F1 score of the best performing model were 0.8321,
0.8302, and 0.8310 respectively. ROC-AUC scores of High risk,
Low, and Medium risk classes were 0.9785, 0.9686, and 0.8943
respectively. The explainable AI analysis was done, and target
leakage was checked. Three results were observed during audit:
no predictor is a fixed function of another predictor; no threshold
value exists for any class; deleting the composite scenario risk
index reduces the accuracy score by 6.66 points. Conclusion
includes a forecast of the changes in risk factors between 2027
and 2030.
