The rapid progress in deep generative models has led to the creation of hyper-realistic synthetic media that easily evades human perception. While modern networks can generate convincing fake assets, existing forensic tools struggle to generalize because they are optimized either for isolated images or continuous video streams, but rarely both. To address this limitation, we propose a unified, cross-medium deep learning framework designed to evaluate authenticity across independent images. The proposed framework evaluates independent images by utilizing a custom Convolutional Neural Network (CNN) baseline alongside advanced Vision Transformers (ViT) and Hierarchical Mixed-Attention (MaxViT) architectures to capture micro-texture spatial anomalies. The framework was comprehensively evaluated on large-scale datasets under real-world data distributions. In the image-level analysis, the baseline CNN achieved an accuracy of 55.74%, which improved to 89.11% with the ViT backbone, and peaked at 93.40% using the MaxViT architecture. Ultimately, this work provides a highly generalizable and robust solution that significantly advances digital media forensics and multimedia verification systems.
