An Explainable Machine Learning Framework for Cyberbullying Detection in Social Media Comments: Hybrid Lexical–Sublexical Modeling with Global and Local Explanations

This research provides an explainable and high-performing cyberbullying detection framework by combining (1) hybrid lexical–sublexical text representation (TF-IDF with word-level and character-level) (2) optimized LinearSVC classification (3) explanation for decisions based on SHAP and LIME. The limitations of existing black-box detection models are overcome as the proposed approach reasoning about the predictions at the feature level in a transparent manner. The framework is evaluated using a custom dataset of 22,630 social media comments, and achieves a score of 97.77% accuracy, 98.04% F1-score, and 0.9976 ROC-AUC, with explanation faithfulness reported using token-deletion analysis. This research is making strides towards more trustworthy AI-assisted content moderation while also providing high predictive performance alongside interpretable decision explanations.