A silent tokenisation failure in the standard Bangla pre
processing pipeline is quantified: 83.2% of vocabulary
types are discarded at a cost of 0.0873 macro-F1, and a
one-line correction recovers them.
• Runs and rank-correlation tests establish that the corpus
is batch-ordered rather than random, and the resulting
evaluation inflation under naive splitting is measured.
• The construct validity of the religious hate label is ques
tioned, since an automatic heuristic detects no religious
marker in 62.7% of comments; a partial-input baseline
and per-target recall make the consequences visible.
• The Bangla T–V pronominal register is identified and
statistically validated as a correlate of the hate label,
under controls that separate a genuine sociolinguistic
signal from a corpus artefact.
• A six-item audit checklist is proposed for Bangla and
other low-resource hate speech resource papers.
