Real-Time Detection of Toxic and Hate Speech in Esports Live Chat: A Comparative Study of Machine Learning and Deep Learning Approaches

The significant research contribution is the creation of the first multi-game annotated esports live-chat dataset for toxicity detection, containing 4,957 consensus-labeled messages from five major games. The study also provides a systematic comparison of eight ML, DL, and Transformer models, showing that DistilBERT achieved the best F1-score (0.8151), while Logistic Regression remained a competitive, much faster option for real-time moderation.