A Vision Transformer-Based Culturally Contextualized Bangla Image Captioning

Vision-Language Models (VLMs) have advanced rapidly for high-resource languages such as English, yet Bengali remains severely underserved in multimodal AI research. This paper presents a culturally grounded hybrid Bengali image captioning system built upon BanglaVision45K, a dataset of 45,000 images paired with 225,000 human-verified Bengali captions
(five per image), assembled from four complementary sources:
Flickr30k (translated via Google Translate), BNature, Bornon,
and a custom collection of 4,000 manually annotated Bangladeshi
cultural images. One-third of the dataset (33.33%, approximately
15,000 images) represents native Bangladeshi cultural contexts.
The corpus contains 46,499 unique Bengali tokens with an
average caption length of 9.53 words. We propose a novel hybrid
architecture combining a frozen SigLIP2 Vision Transformer
encoder (layers 1–8 frozen, layers 9–12 trainable, patch size 32)
with a custom hybrid decoder integrating 4-layer LSTM, 4-layer
GRU, and BanglaGPT, connected through a Luong multiplicative
attention-based fusion mechanism. Through systematic bench-
marking of eight architectures, SigLIP2 + Att-GRU Ex achieves
the best ablation efficiency-performance trade-off (BLEU-1:
0.5360, METEOR: 0.3821, ROUGE-L: 0.4631, CIDEr: 0.2205).
Our proposed final hybrid model (SigLIP2 + LSTM + GRU +
BanglaGPT) trained on the full BanglaVision45K (90/5/5 split)
achieves state-of-the-art performance: BLEU-1: 0.8356, BLEU-
2: 0.7219, BLEU-3: 0.6119, BLEU-4: 0.5194, METEOR: 0.7004,
ROUGE-L: 0.7414, CIDEr: 0.5287, outperforming all baselines
by over 15% across all metrics.