Quantifying Social Bias in LLMs for Bangla: A Likelihood-Based Comparative Analysis

This study provides a comparative evaluation of stereotypical social bias in three instruction-tuned LLMs on the Bangla BanStereoSet benchmark across nine bias categories. A key contribution is the systematic comparison of whole-sentence and candidate-token-conditioned likelihood scoring, demonstrating that the scoring methodology can substantially alter measured bias. The findings further show that Bangla-language specialization alone does not guarantee reduced stereotypical bias, highlighting the need for language- and script-aware fairness evaluation for Bangla LLMs.