This work contributes a lightweight, training-free method for detecting hallucinations in large language models that requires no model retraining, no access to model internals, and no specialized hardware, running on a single consumer GPU. By decomposing self-consistency into token, entity, and semantic granularities and fusing them with interpretable per-model weights, it delivers competitive detection while avoiding the computational and energy cost of gradient-based or white-box approaches. We further show that the most effective signal varies with a model’s tuning regime, and we release the full evaluation pipeline to support reproducible, resource-efficient hallucination research.
