Authors: Research Scholar Priya Mishra, Dr. Deepika Pathak
Abstract: The rapid evolution of deepfake generation technologies has created significant challenges for digital media authentication, cybersecurity and misinformation prevention. Modern deepfake images generated using advanced artificial intelligence techniques exhibit highly realistic visual characteristics, making manual identification increasingly difficult. This paper presents a comprehensive review of efficient hybrid CNN-transformer models for benchmark-driven image deepfake detection and analysis. The study critically examines recent advances in Convolutional Neural Networks (CNNs), transformer-based architectures, hybrid feature fusion frameworks and explainable AI techniques for identifying manipulated images, and analyses benchmark-driven evaluation strategies, multimodal feature representation, manipulation localisation and attention-based contextual learning. Twelve studies published in 2025 are reviewed and organised into three themes covering hybrid architectures, benchmark-driven feature analysis, and cybersecurity challenges. The reviewed evidence is consolidated into comparative tables that map each study to its focus, approach, principal finding and limitation, alongside a component-level comparison of CNN, transformer and hybrid designs and an assessment of modality coverage across the corpus. Major challenges including adversarial attacks, computational complexity, cross-domain generalisation, multimodal deepfakes and benchmark inconsistency are discussed, and seven research gaps are identified and mapped to corresponding future research directions focusing on explainable AI integration, multimodal fusion, computational optimisation and generalised feature learning.
International Journal of Science, Engineering and Technology