A Benchmark of Fake Review Detection Using Machine Learning in E-Commerce

The increasing use of online reviews has made reliable review information essential for consumer decision-making, while deceptive reviews reduce trust in e-commerce platforms. This study addresses fake review detection in Banglish, a code-mixed form of Bangla and English for which benchmark datasets and specialized detection frameworks are limited. A dataset of 25,588 reviews was developed and annotated into fake and not_fake classes, capturing Bangla script, Romanized Bangla, English, transliteration, and informal code-mixing. A hybrid BanglaBERT-BiLSTM model is proposed, where BanglaBERT learns contextual semantic representations and BiLSTM captures sequential dependencies. The proposed system achieved 96.82% accuracy, 96.83% precision, 96.82% recall, 96.44% F1-score, and 99.47% ROC-AUC on the unseen test set. It outperformed the evaluated traditional and transformer-based baselines. LIME and SHAP analyses further provide interpretable evidence about the textual features influencing predictions. The dataset and benchmark provide a foundation for future research in Banglish NLP and deceptive review detection.