This work proposes SAFER-RAG, a lightweight and interpretable hallucination detection framework that integrates lexical/factual, semantic, and NLI-based evidence—claim features with confidence calibration, selective prediction, and TreeSHAP explainability. Beyond in-domain detection, the study evaluates reliability under cross-task, cross-generator, source-domain, and external benchmark shifts. The results show that discrimination can remain strong even when calibration degrades, highlighting the need to assess confidence reliability separately. On RAGTruth, Platt scaling reduced ECE from 0.0920 to 0.0583, while abstaining on the least-confident 20% of predictions reduced error risk by 23.6%.
