Seeing Beyond the Fake: Detecting Deepfakes Using Deep Learning-Based Computer Vision

The main advantage of this research is the creation of the lightweight, leakage-aware spatial-temporal framework for deepfake detection. The work proposes a custom CNN with channel-wise attention, batch normalization, dropout and L_2 regularization, with approximately 511K trainable parameters. It also introduces frame-level detection to video-level analysis with an Attention-BiLSTM and fusion-based classification approach. One methodological contribution is the video-first data splitting to avoid data leakage by keeping all of the frames from the same video in either the training set or the test set. The study also includes a comparison of the spatial, temporal and fusion approaches, where the fusion approach attained 93.50% accuracy at the video level and 97.99% AUC.