Machine Learning for Literacy Rate Prediction in Bangladesh: Leveraging Engineered Socio-Economic Features and Explainable AI

This study presents a machine learning framework for predicting and analyzing literacy rates in Bangladesh, using a wide range of socio-economic and demographic data. The goal is to provide a practical statistical tool that helps explain the complex factors shaping national literacy. To gather useful information from the data such as GPD per capita and population density, we applied data preprocessing and feature engineering. We tested several regression models and among them the Random Forest regression model performed best achieving high prediction accuracy with a $R^{2}$ score of 0.992. A major contribution of this study is the use of Explainable AI (XAI) with SHapley Additive exPlanations (SHAP). This allowed us not only to predict literacy rates but also to understand how different factors influence the model’s results. The SHAP analysis showed that population patterns and economic indicators such as population density and gross national income (GNI) are the most important drivers of literacy outcomes. In summary, this study offers a highly accurate predictive model and also explains the complex relationships between different variables making it a valuable resource for both researchers and policymakers.