This study introduces a fine-grained Bangla toxic comment classification framework that distinguishes Explicitly Toxic, Subtly Toxic and Neutral content. It provides a controlled comparison of six models across binary and three-class settings using a unified dataset of 20,116 manually labeled Bangla comments, revealing a consistent 29–39 F1-point performance drop when subtle toxicity is introduced. The findings highlight subtle toxicity as a key unresolved challenge in Bangla NLP.
