<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.3" xml:lang="en">
  <front>
    <journal-meta> 
	   <journal-id journal-id-type="publisher-id">tijfs</journal-id>
	   <journal-title-group> 
	    <journal-title>The International Journal of Frontier Sciences</journal-title> 
		<abbrev-journal-title abbrev-type="publisher">Int. J.  Front. Sci.</abbrev-journal-title>
		<abbrev-journal-title abbrev-type="pubmed">The International Journal of Frontier Sciences</abbrev-journal-title> 
	  </journal-title-group>
	 <issn pub-type="epub">2618-0367</issn> 
	 <publisher>
	    <publisher-name>FRONTIER SCIENCE ASSOCIATES</publisher-name> 
	 </publisher>
	</journal-meta> 
    <article-meta>
      <article-id pub-id-type="doi">10.37978/tijfs.v08i01.004</article-id>
      <article-id pub-id-type="publisher-id">tijfs-8-4</article-id>
      <article-categories>
        <subj-group>
          <subject>Article</subject>
        </subj-group>
      </article-categories>
      <title-group>
        <article-title>Toward Efficient Fake Frame Detection in Video Using Deep Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Alam</surname>
            <given-names>Iftikhar</given-names>
          </name>
          <xref rid="af1-tijfs-8-4" ref-type="aff">1</xref>
          <xref rid="c1-tijfs-8-4" ref-type="corresp">*</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Ahsan Kamran</surname>
            <given-names>Malik</given-names>
          </name>
          <xref rid="af1-tijfs-8-4" ref-type="aff">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Ullah</surname>
            <given-names>Ehaab</given-names>
          </name>
          <xref rid="af1-tijfs-8-4" ref-type="aff">1</xref>
        </contrib>
      </contrib-group>
      <contrib-group>
        <contrib contrib-type="editor">
          <name>
            <surname>Zar</surname>
            <given-names>Mian Sahib</given-names>
          </name>
          <role>Academic Editor</role>
        </contrib>
        <contrib contrib-type="editor">
          <name>
            <surname>Akhtar</surname>
            <given-names>Muhammad Shoaib</given-names>
          </name>
          <role>Academic Editor</role>
        </contrib>
      </contrib-group>
      <aff id="af1-tijfs-8-4"><label>1</label>Department of Computer Science, City University of Science and Information Technology, Peshawar 25000, Pakistan</aff>
      <author-notes>
        <corresp id="c1-tijfs-8-4"><label>*</label>Correspondence: iftikharalam@<email>cusit@edu.pk</email></corresp>
      </author-notes>
      <pub-date publication-format="electronic" date-type="pub" iso-8601-date="2026-08-25">
        <day>25</day>
        <month>08</month>
        <year>2026</year>
      </pub-date>
      <volume>8</volume>
      <issue>1</issue>
      <elocation-id>4</elocation-id>
      <history>
        <date date-type="received">
          <day>06</day>
          <month>02</month>
          <year>2026</year>
        </date>
        <date date-type="accepted">
          <day>10</day>
          <month>08</month>
          <year>2026</year>
        </date>
      </history>
      <permissions>
        <copyright-statement>&#xA9; 2026 by the TIJFS.</copyright-statement>
        <copyright-year>2026</copyright-year>
        <license license-type="open-access">
          <license-p>This is an open-access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (<ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link>).</license-p>
        </license>
      </permissions>
      <abstract>
        <p><bold>Background:</bold> Deepfake technology is a major social concern due to the rapid development of Artificial Intelligence (AI), especially in machine learning (ML) and deep learning (DL). Deepfakes are artificially modified videos and images that effectively change a person&#x2019;s facial features or expressions to misrepresent reality. These videos and images, often unnoticed by casual observers, present significant ethical, political, and social implications, as they can spread misinformation, damage reputations, and influence public perception. This study is an attempt to detect artificially modified videos by analyzing each frame using DL. We use ResNet-50 architecture, a well-known Convolutional Neural Network (CNN) model, to identify tampered videos. <bold>Methods:</bold> The system is trained using the Celebrity Deep Fake dataset, which includes numerous original and fake video samples. The model assesses whether each frame is original or tampered with after the videos have been split into frames. The system is tested and evaluated using standard metrics, including accuracy, precision, recall, and F1-score. <bold>Results:</bold> The model achieved 82.33% accuracy, 76.70% precision, 89.15% recall, and an F1-score of 82.46%. These results indicate that deepfake videos were correctly detected and that the model was efficient at identifying most real deepfake instances. In addition, the F1-score of 82.46% is further evidence of the model&#x2019;s stability, as it ensures both high accuracy and coherence across numerous cases. <bold>Conclusions:</bold> The findings indicate that the proposed model works effectively compared to existing deepfake methods. In the future, we intend to use a richer dataset with more resources, which may further enhance the accuracy of the model.</p>
      </abstract>
      <kwd-group>
        <kwd>deepfake videos</kwd>
        <kwd>deep learning model</kwd>
        <kwd>ResNet-50</kwd>
        <kwd>video frames detection</kwd>
        <kwd>video forensics</kwd>
      </kwd-group>
    <custom-meta-group>
        <custom-meta>
          <meta-name>How to cite</meta-name>
          <meta-value>Alam, I.; Ahsan Kamran, M.; Ullah, E. Toward Efficient Fake Frames Detection in Video Using Deep Learning. <italic>TIJFS</italic>&#xA0;<bold>2026</bold>, <italic>8</italic>, 4. DOI:<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.37978/tijfs.v08i01.004">10.37978/tijfs.v08i01.004</ext-link>.</meta-value>
        </custom-meta>
      </custom-meta-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec1-tijfs-8-4" sec-type="intro">
      <title>1. Introduction</title>
      <p>In today&#x2019;s rapidly evolving digital content ecosystem, the internet is flooded with millions of videos and images each day [<xref ref-type="bibr" rid="B1-tijfs-8-4">1</xref>], some of which are tampered with using deepfake techniques. Deepfakes are fake/altered videos that are created with the help of AI tools and deep learning (DL) techniques [<xref ref-type="bibr" rid="B2-tijfs-8-4">2</xref>,<xref ref-type="bibr" rid="B3-tijfs-8-4">3</xref>]. These technologies can realistically recreate a person&#x2019;s face and expression so that it becomes hard to judge whether the content is real or fake. This represents a major problem in every field, including politics, journalism, and social media. [<xref ref-type="bibr" rid="B4-tijfs-8-4">4</xref>]. For example, a fake video can be generated using AI tools to make it appear as if a politician said something that they never said, or to show a celebrity performing something that they never did. The video will quickly become viral due to the confusion, spreading false information and damaging public opinion. One of the major concerns regarding fake videos is that they destroy the credibility of social media and public trust [<xref ref-type="bibr" rid="B5-tijfs-8-4">5</xref>].</p>
      <p>The loss of trust in social media platforms can have a serious effect on society, including political instability, online harassment, financial fraud, and blackmailing. Deepfakes are becoming more advanced day by day, and therefore, detection becomes more difficult. There is a strong need for an automatic system that can detect these fake videos quickly and accurately. This study focuses on such a system, which builds on DL techniques. The purpose is to develop a model for a tool that analyzes videos, helping individuals, along with organizations and media outlets, identify fake content before it spreads. There is a need for effective techniques to identify and deal with these manipulated videos.</p>
      <p>With the massive growth of knowledge regarding deepfakes, researchers have created numerous algorithms and tools that are in development [<xref ref-type="bibr" rid="B6-tijfs-8-4">6</xref>,<xref ref-type="bibr" rid="B7-tijfs-8-4">7</xref>]. These tools are useful for identifying deepfakes, but they have certain gray areas that limit their usefulness. There are some web applications available, such as Deep Fake Detector, Detect Deepfakes, and Alterations. These tools help us to detect whether uploaded videos/movies are real or fake.</p>
      <p>While these tools do help in some cases, they nonetheless fail to address important issues like ethical concerns, legal challenges, and public awareness. The above factors are very important when it comes to dealing with the harmful effects of fake videos on society. DeepFakeDetector (<uri>https://sightengine.com/detect-deepfakes</uri>) has a major limitation, which is that it only deals with analyzing videos of up to 5 megabytes (MB) in size. This is a major issue because in the modern era of digital content, videos are larger in size and of higher quality. As a result, this tool does not fulfill the requirements of the currently uploaded content shared on social platforms. Deepfake videos are becoming more common and more realistic, which makes it difficult to identify whether videos are original or fake [<xref ref-type="bibr" rid="B8-tijfs-8-4">8</xref>]. Numerous researchers around the world have proposed various techniques and models to detect fake videos [<xref ref-type="bibr" rid="B9-tijfs-8-4">9</xref>].</p>
      <p>The following few paragraphs discuss how deepfake detection models work, what kind of datasets they use, how accurate they are, and what limitations they have. This section also provides detail information on the datasets that these researchers have used, such as FaceForensics++, DFDC, and Celeb-DF [<xref ref-type="bibr" rid="B10-tijfs-8-4">10</xref>]. Datasets have been used widely in the field and contain both real and fake videos for training and testing models.</p>
      <p>Classification plays an important role in deepfake image/video detection [<xref ref-type="bibr" rid="B11-tijfs-8-4">11</xref>,<xref ref-type="bibr" rid="B12-tijfs-8-4">12</xref>]. After the extraction of features from the video (like facial movements, eye blinking, head motion, or mouth movement), we need a model that can classify them into one of two categories: real or fake. Most researchers prefer to use DL models for this classification task due to their deep learning capabilities. The deep Convolutional Neural Network model, which is used for image classification, is known for its high number of layers and its architecture. Numerous researchers have combined both CNN and Long Short-Term Memory (LSTM) to take advantage of both image-based and time-based features. The CNN extracts features from each of the frames, and the LSTM analyzes the sequence of these features, in order to make the final decision at the end.</p>
      <p>In addition to these approaches, some studies use transfer learning with pre-trained models like the Visual Geometric Group-16 (VGG-16), Residual Network 50 (ResNet-50), LSTM, Residual Networks with Aggregated Transformations, and EfficientNet (ResNeXt) [<xref ref-type="bibr" rid="B13-tijfs-8-4">13</xref>]. These models are pre-trained on large image datasets, and they can quickly learn new tasks with less training data. Using the technique of transfer learning has sped up the development and improved the accuracy of these models. Although these DL techniques are powerful and have achieved high accuracy, they also present some challenges. They usually require high-performance hardware such as a Graphics Processing Unit or Tensor Processing Unit for the processing of large amounts of data. Also, running them in real-time environments, such as live video detection, can be slow or inefficient [<xref ref-type="bibr" rid="B14-tijfs-8-4">14</xref>,<xref ref-type="bibr" rid="B15-tijfs-8-4">15</xref>]. In conclusion, classification techniques are always the key part of any deepfake detection system.</p>
      <p>Aarti Karandikar et al. [<xref ref-type="bibr" rid="B16-tijfs-8-4">16</xref>] proposed a deepfake detection method that makes use of a transfer learning-enhanced VGG-16 CNN. This technique finds visual irregularities, especially between facial features, in order to identify deepfakes. This system had trouble with low-resolution deepfakes and was sensitive to compression artifacts. Similarly, Hanqing Zhao et al. [<xref ref-type="bibr" rid="B17-tijfs-8-4">17</xref>] presented a Multilayer Feedforward Neural Network that can record artifacts at different levels of granularity. The model performed well, particularly when it came to identifying low-quality deepfakes. However, the model&#x2019;s high energy and computing costs made real-time adoption challenging. TensorFlow and Keras were used to assess the approach on the FaceForensics++ dataset.</p>
      <p>Yash Doke et al. [<xref ref-type="bibr" rid="B18-tijfs-8-4">18</xref>] created a hybrid model that uses LSTM to capture temporal discrepancies and ResNetCNN to extract features. By concentrating on facial and kinematic abnormalities, this approach produced high detection accuracy and real-time performance. Its significant processing burden during frame-level analysis was its primary flaw. The model was developed using Python frameworks and trained on FaceForensics++ with a 70% training set and 30% test set. Alakananda Mitra et al. [<xref ref-type="bibr" rid="B19-tijfs-8-4">19</xref>] created a deepfake detection model using a CNN combined with a classifier. Despite achieving high precision scores, the model was not resistant to different video quality and compression levels. It was trained using the FaceForensics++ dataset and implemented with Keras and TensorFlow.</p>
      <p>Nicolo Bonettini et al. [<xref ref-type="bibr" rid="B20-tijfs-8-4">20</xref>] suggested a hybrid approach to identify facial changes in deepfake films by fusing DL methods with conventional computer graphics. Although the CNN-based system was quite accurate, its real-time application was limited by modest processing delays. The FaceForensics++ and Deepfake Detection Challenge (DFDC) datasets in PyTorch were also used to assess the model. Ankur Nagulwar et al. [<xref ref-type="bibr" rid="B21-tijfs-8-4">21</xref>] used Multi-task Cascaded Convolutional Networks (MTCNNs) to identify irregularities in noise and face features in video frames. The system successfully authenticated videos with single faces and achieved 70% accuracy. But when there were several faces in a single shot, its performance degraded. The FaceForensics++ dataset, which consists of films that have been transformed into frames for analysis, was used for the evaluation. Nimitt Patel et al. [<xref ref-type="bibr" rid="B22-tijfs-8-4">22</xref>] used an LSTM and ResNext CNN to identify video frames in a deepfake detection model. However, the model&#x2019;s usefulness was limited due to the limited visual content examined, and it did not analyze audio. The preprocessing with cropping and resizing, FaceForensics++, was used for training.</p>
      <p>Neeraj Guhagarkar et al. [<xref ref-type="bibr" rid="B23-tijfs-8-4">23</xref>] provided a thorough analysis of several AI and machine learning-based deepfake detection techniques. They examined temporal and spatial approaches and highlighted high-performing designs, such as CNN+LSTM. The study identified issues with low-quality movies and real-time identification despite the extensive use of datasets such as Celeb-DF, University of Albany Deep Fake Video (UADFV), and Reddit Deepfakes. TensorFlow, Xception, and Visual Geometry Group (VGG) were commonly used tools. Ahmed Hatem Soudy et al. [<xref ref-type="bibr" rid="B24-tijfs-8-4">24</xref>] suggested a hybrid deepfake detection system that uses a majority voting method and combines CNNs with vision transformers. The eyes, nose, and entire face were the focal points of their model. Although it needed a lot of processing power and might have missed face alterations in other regions, it was able to reach up to 80% accuracy. The FaceForensics++ and DFDC datasets were used to train and assess the system.</p>
      <p>Similarly, Arash Heidari et al. [<xref ref-type="bibr" rid="B25-tijfs-8-4">25</xref>] carried out a study of the literature on DL-based deepfake detection techniques. CNNs, GANs, RCNNs, and other models for image, video, and audio detection were discussed in the study. Methods outside of DL were not included; they were grouped according to performance and application. They concentrated on models built using platforms like TensorFlow and Keras and compared them using datasets including CelebA, TIMIT, and DFDC. A detailed yet comprehensive literature review is presented in <xref ref-type="table" rid="tijfs-8-4-t001">Table 1</xref>.</p>
	  <table-wrap id="tijfs-8-4-t001" position="anchor">
        <object-id pub-id-type="pii">tijfs-8-4-t001_Table 1</object-id>
        <label>Table 1</label>
        <caption>
          <p>Most relevant studies about fake frame detection in terms of their datasets, methodologies, and experimental environments</p>
        </caption>
        <table>
          <thead>
            <tr>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Studies</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Dataset</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Methodology </th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Simulation Environment</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Conclusion</th>
            </tr>
          </thead>
          <tbody>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin"><bold>Ju, Hu et al. [<xref ref-type="bibr" rid="B26-tijfs-8-4">26</xref>]</bold></td>
              <td align="left" valign="middle" style="border-bottom:solid thin">DFDC,<break/>Celeb-DF, DFD, FaceForensics++<break/>(with demographic annotations)</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">DAG-FDD<break/>(agnostic) and DAW-FDD<break/>(aware) models using CVaR loss to ensure fairness in deepfake detection</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">PyTorch; Dlib reprocessing; NVIDIA RTX A6000; frame size<break/>380 &#xD7; 380; tested on ResNet50, Xception</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Fair detection without losing accuracy</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin"><bold>Pei, Zhang et al. [<xref ref-type="bibr" rid="B27-tijfs-8-4">27</xref>]</bold></td>
              <td align="left" valign="middle" style="border-bottom:solid thin">FaceForensics++, CelebDF, DFD, DFDC</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Categorize detection techniques and propose OOD benchmarks</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Literature-based evaluation without model testing</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Cross-domain fairness without accuracy loss</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin"><bold>Taeb and Chi [<xref ref-type="bibr" rid="B28-tijfs-8-4">28</xref>]</bold></td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Kaggle Deepfake Detection<break/>Challenge<break/>(subset)</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Used CNN-<break/>RNN, 2D/3D CNN, Two-<break/>Stream with optical flow and F1/BCE loss</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Google Colab with OpenCV and TensorFlow</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Achieved 79% accuracy without transformers</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin"><bold>Passos, Jodas et al. [<xref ref-type="bibr" rid="B29-tijfs-8-4">29</xref>]</bold></td>
              <td align="left" valign="middle" style="border-bottom:solid thin">FaceForensics++, DFDC,<break/>Celeb-DF,<break/>WildDeepfake,<break/>UADFV</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Reviewed 100+ studies using CNNs, RNNs, GANs, and Transformers; categorized<break/>methods into spatial, temporal, and hybrid approaches</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Focused mainly on facial manipulations; lacked cross-modal and audio analysis</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Analytical review without implementation; highlights dataset uses and model<break/>categorization</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin"><bold>Rajagukguk, Kencana et al. [<xref ref-type="bibr" rid="B30-tijfs-8-4">30</xref>]</bold></td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Kaggle (589 real, 700 fake)</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Used ResNet50 with transfer learning for binary face classification</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Achieved 76% training and 53% test accuracy, indicating overfitting</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Trained in TensorFlow for 30 epochs using data Augmentation and a frozen base model</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin"><bold>Orlando and Al Rivan [<xref ref-type="bibr" rid="B31-tijfs-8-4">31</xref>]</bold></td>
              <td align="left" valign="middle" style="border-bottom:solid thin">ISIC 2019 dataset with 25,331 images of 8 skin cancer types</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Used CNN models (LeNet and VGG-16) for classification</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Implemented in Python using<break/>Adam optimizer, data augmentation, batch size 32, 30 epochs</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">VGG-16 achieved 73.22% accuracy (better than LeNet) but required more training time (354 s per epoch)</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin"><bold>Notshe, Xiao et al. [<xref ref-type="bibr" rid="B32-tijfs-8-4">32</xref>]</bold></td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Subset of Kaggle Deepfake Detection<break/>Challenge (small and<break/>imbalanced)</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Tested 2D<break/>CNN, 3D CNN,<break/>CNN-RNN, and Two-Stream<break/>CNNs using both spatial and temporal features</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Implemented in Google Colab with TensorFlow and OpenCV;<break/>used optical flow</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Temporal models (e.g., CNN-RNN) outperformed<break/>others (up to 79%); no transformers tested</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin"><bold>Samuel [<xref ref-type="bibr" rid="B33-tijfs-8-4">33</xref>]</bold></td>
              <td align="left" valign="middle" style="border-bottom:solid thin">ImageCLEF med, OASISMRI, TCIA-CT, and NEMA-CT<break/>datasets</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Fused CNN using shallow and deep layers; applied transfer learning on selected layers</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Used Caffe framework, Intel i7 (16 GB RAM), LR 1e-6, momentum 0.9</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Achieved better classification; limited by model complexity and small datasets</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin"><bold>Alzahrani and Rawat [<xref ref-type="bibr" rid="B34-tijfs-8-4">34</xref>]</bold></td>
              <td align="left" valign="middle" style="border-bottom:solid thin">FaceForensics++ (Face2Face subset)</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Used CNNs with optical flow input to detect motion irregularities in deepfake videos</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Implemented<break/>with PWC-Net and TensorFlow.<break/>RGB optical flow processed via VGG16/Res Net50</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">High accuracy on Face2Face forgeries; limited by dependence on clean optical flow data</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin"><bold>Bankar, Pawar et al. [<xref ref-type="bibr" rid="B35-tijfs-8-4">35</xref>]</bold></td>
              <td align="left" valign="middle" style="border-bottom:solid thin">FaceForensics++<break/>(NeuralTextures,<break/>FaceSwap,<break/>DeepFakes,<break/>Face2Face)</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Benchmarked deepfake detector using over 1.8Mmanipulatedimages; used XceptionNet for classification</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Tested across compression levels; used Face2Face for tracking and frame-level classification</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Detection accuracy drops with<break/>compression; emphasizes the need for diverse datasets and standard benchmarks</td>
            </tr>
          </tbody>
        </table>
      </table-wrap>
      <p>When a fake video, which is generated very carefully and with advanced tools, looks very realistic, these tools are unable to detect the changes made in the video properly. This also leads to inaccurate results, especially when it comes to detecting fake videos that are professionally made/edited. The limitations clearly show the need for a more advanced and more reliable fake video detection system, which can not only process large and high-quality videos but can also provide accurate and trustworthy results. The system should be user-friendly and should be capable of helping media agencies and legal authorities to identify fake content effectively. This study aims to fill the gaps in the research by developing a more efficient and powerful fake video detection tool to overcome these challenges.</p>
      <p>The accuracy and dependability of current detection techniques are sometimes insufficient to handle the increasing difficulty of identifying manipulated material, especially deepfakes. Due to this restriction, it is challenging to identify small alterations in video content, which enables bogus videos to aid in the dissemination of false information. The social, political, and media landscapes are at serious risk, and public confidence in digital media is eroded as a result.</p>
      <p>The following are the aims and objectives of the study:</p>
	  <list list-type="bullet">
        <list-item>
          <p>To review the state of the art of the literature in detecting deepfake videos.</p>
        </list-item>
        <list-item>
          <p>To propose a DL approach for detecting fake videos.</p>
        </list-item>
        <list-item>
          <p>To build a web application as an application prototype.</p>
        </list-item>
      </list>
      <p>This study addresses the issues and challenges that offer reliable and user-friendly problem solutions for verifying whether a video is authentic or not. The system is built to adapt to all these limitations and provide accurate and high-quality results. By using the advanced DL model and user-friendly web interface, the tool can be used for detecting deepfake videos. In addition, we also compared the accuracy of each method to understand how well it performs. Some of the models must achieve high accuracy but may be limited to a specific type of fake video, or they might not perform as well as they are expected to in real-world scenarios where lighting and quality vary. Through this research, by reviewing all these models, datasets, and results, we gain a better understanding of what works perfectly and what does not.</p>
    </sec>
    <sec id="sec2-tijfs-8-4">
      <title>2. Materials and Methods</title>
      <p>This study focuses on detecting fake videos, also known as deepfakes, using DL techniques. The methodology includes several key steps: dataset selection, preprocessing, model training, and evaluation, as shown in <xref ref-type="fig" rid="tijfs-8-4-f001">Figure 1</xref>. Each of these steps plays a vital role in ensuring that the model accurately distinguishes real from fake videos. The proposed system is designed for the detection of fake videos using a DL-based approach.</p>
	  <fig id="tijfs-8-4-f001" position="anchor">
        <label>Figure 1</label>
        <caption>
          <p>Proposed steps of overall workflow, including dataset selection, preprocessing, model training, and evaluation. These stages are performed sequentially to develop and assess the proposed model.</p>
        </caption>
        <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="image001.png"/>
      </fig>
      <p>The architecture shown in <xref ref-type="fig" rid="tijfs-8-4-f002">Figure 2</xref> includes several components, such as the selection of the dataset, preprocessing of the dataset, selection of the model, model training, video analysis, and the display of the results on the UI. The DL model is trained and evaluated for fake video detection.</p>
	  <fig id="tijfs-8-4-f002" position="anchor">
        <label>Figure 2</label>
        <caption>
          <p>Proposed architecture. This diagram outlines the training and prediction pipeline of a CNN-based deepfake detection system. During training, real and fake video datasets are split, preprocessed into resized frames, and used to train and export a detection model. During prediction, an uploaded video undergoes the same preprocessing steps before being classified as real or fake by the trained model.</p>
        </caption>
        <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="image002.png"/>
      </fig>
      <sec id="sec2dot1-tijfs-8-4">
        <title>2.1. Dataset Selection and Preprocessing</title>
        <p>We use the Celeb-DF dataset, which is one of the most well-known and challenging deepfake datasets available for academic and research purposes in deepfakes. The dataset was first introduced in ref. [<xref ref-type="bibr" rid="B36-tijfs-8-4">36</xref>]. The main goal of creating this dataset was to address the limitations of earlier deepfake datasets such as University of Albany Deep Fake Video (UADFV) and Faceforensics++, which have noticeable visual artifacts that make the fake content easier to detect.</p>
        <p>This dataset contains 408 real videos that were collected from interviews of different celebrities on YouTube, covering both males and females of different ethnicities, ages, and speaking styles. Moreover, this dataset also contains 795 fake videos that were created using advanced deepfake generation techniques, which focus on making the fake videos look more realistic by preserving head poses, lip sync, and facial expression.</p>
        <p>The DL model works better with images rather than videos. For this purpose, each video, whether it is real or fake, is first split into individual frames using OpenCV, which is a Python library. Each video is processed frame by frame. Moreover, the proposed frames must be in a proper sequence. The frame rate is kept consistent by extracting eight frames per second from 30 frames per second to avoid redundancy and to reduce processing time. These extracted frames serve as input data samples for model detection.</p>
        <p>The images are cropped and resized to a size of 224 &#xD7; 224. This is necessary because the ResNet50 model expects the input of images to be of a fixed dimension. OpenCV is utilized to convert color formats (such as RGB), resize photos to the necessary input shape (224 &#xD7; 224 pixels), and carry out simple tasks like cropping or grayscale conversion when necessary. For DL training, OpenCV makes the video-to-frame conversion dependable and effective. The pixel values of the images are also normalized to a scale of between 0 and 1 to reduce the computational load and to help the model train more efficiently. Each image is assigned a label; Label 0 is given to real images, and Label 1 is assigned to fake images.</p>
      </sec>
      <sec id="sec2dot2-tijfs-8-4">
        <title>2.2. Model Selection</title>
        <p>To achieve the accuracy and reliability of the model, we have selected the ResNet50 model, which is a powerful and widely used CNN model. This model was first introduced by Microsoft Research in ref. [<xref ref-type="bibr" rid="B37-tijfs-8-4">37</xref>].</p>
        <sec id="sec2dot2dot1-tijfs-8-4">
          <title>2.2.1. Comparison with Other CNN Models</title>
          <p>The ResNet50 model is a better, more accurate, and more reliable CNN model than other available models. It extracts features and is a lightweight model compared to the VGG16 model. Also, fine-tuning is easily comparable to InceptionV3 and is more accurate compared to MobileNet. A detailed comparison of these models is provided in <xref ref-type="table" rid="tijfs-8-4-t002">Table 2</xref>.</p>
		  <table-wrap id="tijfs-8-4-t002" position="anchor">
        <object-id pub-id-type="pii">tijfs-8-4-t002_Table 2</object-id>
        <label>Table 2</label>
        <caption>
          <p>Comparative analysis of commonly used deep learning models for deepfake detection</p>
        </caption>
        <table>
          <thead>
            <tr>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Model</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Depth</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Parameters</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Accuracy<break/>(General)</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Suitability for Deepfake<break/>Detection</th>
            </tr>
          </thead>
          <tbody>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>VGG16</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">16 Layers</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">138M</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Good</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Large model, slower training.</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>VGG19</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">19 Layers</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">144M</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Good</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">A very large model, limited to image classification.</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>Inception V3</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">48 layers</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">24M</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Excellent</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Complex architecture, harder to fine-tune.</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>MobileNet V2</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">53 Layers</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">3.5M</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Fair</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Lightweight but less accurate.</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>EfficientNet B0</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">82 Layers</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">5.3M</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">very Good</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Efficient, but might underperform on complex tasks.</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>ResNet50</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">50 Layers</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">23M</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Excellent</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Balanced: good depth, accuracy, and speed.</td>
            </tr>
          </tbody>
        </table>
      </table-wrap>
        </sec>
        <sec id="sec2dot2dot2-tijfs-8-4">
          <title>2.2.2. Training Process</title>
          <p>The training set used to train the model contains 70% of the data from the dataset. The testing set is used to evaluate the model after training, and contains 15% of the data from the dataset, and the validation set is used to check how well the model is learning during the training session. It also contains 15% of the data from the dataset.</p>
          <p>During training, the model uses the training set, which is already preprocessed with real and fake video frames. The actual goal is to enable the model to properly learn the patterns and features that distinguish the real frames from the fake frames. The Adam optimizer is used for updating the model weights during training. Adam is widely used because it adapts the learning rate during training and combines the benefits of two other optimizers, AdaGrad and RMSProp. It is efficient, fast, and works well for large datasets, as well as for DL models like ResNet50.</p>
          <p>TensorFlow manages the training loop, automatically computes gradients, and updates model weights to minimize loss over time. Keras, with its simple syntax, makes the process of defining models, compiling them, and evaluating their performance more intuitive for quick experimentation. These tools are crucial for implementing and testing different DL techniques during the fake video detection pipeline. At a predetermined frame rate (e.g., eight frames per second), each video is divided into individual image frames, which the model then processes separately.</p>
          <p>The model is trained for 27 epochs. The early stopping technique monitors the validation loss. When the loss does not improve for a certain number of epochs, the training is automatically stopped. This provides the benefit of avoiding training that might harm the model&#x2019;s ability to generalize. The model checkpoints save the best version of the model, depending on the validation performed automatically. When training is completed, the final model will be saved as Hierarchical Data Format Version 5 (HDF5), which is also popularly known as the .h5 format. This model stores the architecture of the model, weights, and the optimized state, which makes it very easy to reuse the model for prediction and deployment in different web applications.</p>
        </sec>
      </sec>
      <sec id="sec2dot3-tijfs-8-4">
        <title>2.3. Evaluation Metrics</title>
        <p>The performance of deepfake detection was evaluated using the following standard classification matrices [<xref ref-type="bibr" rid="B38-tijfs-8-4">38</xref>].</p>
        <sec id="sec2dot3dot1-tijfs-8-4">
          <title>2.3.1. Accuracy</title>
          <p>Accuracy measures how well the model classifies videos as real or fake. It is calculated as the ratio of correctly classified videos (both real and fake) to the total number of videos. Our system achieved 82.33% accuracy, computed as</p>
		  <disp-formula id="FD1-tijfs-8-4">
            <label>(1)</label>
            <mml:math id="mm1" display="block">
              <mml:semantics>
                <mml:mrow>
                  <mml:mi>A</mml:mi>
                  <mml:mi>c</mml:mi>
                  <mml:mi>c</mml:mi>
                  <mml:mi>u</mml:mi>
                  <mml:mi>r</mml:mi>
                  <mml:mi>a</mml:mi>
                  <mml:mi>c</mml:mi>
                  <mml:mi>y</mml:mi>
                  <mml:mo>=</mml:mo>
                  <mml:mstyle scriptlevel="0" displaystyle="true">
                    <mml:mfrac>
                      <mml:mrow>
                        <mml:mi>T</mml:mi>
                        <mml:mi>P</mml:mi>
                        <mml:mo>+</mml:mo>
                        <mml:mi>T</mml:mi>
                        <mml:mi>N</mml:mi>
                      </mml:mrow>
                      <mml:mrow>
                        <mml:mi>T</mml:mi>
                        <mml:mi>P</mml:mi>
                        <mml:mo>+</mml:mo>
                        <mml:mi>T</mml:mi>
                        <mml:mi>N</mml:mi>
                        <mml:mo>+</mml:mo>
                        <mml:mi>F</mml:mi>
                        <mml:mi>P</mml:mi>
                        <mml:mo>+</mml:mo>
                        <mml:mi>F</mml:mi>
                        <mml:mi>N</mml:mi>
                      </mml:mrow>
                    </mml:mfrac>
                  </mml:mstyle>
                </mml:mrow>
              </mml:semantics>
            </mml:math>
          </disp-formula>
        </sec>
        <sec id="sec2dot3dot2-tijfs-8-4">
          <title>2.3.2. Precision</title>
          <p>Precision measures a model&#x2019;s dependability in identifying a fake video [<xref ref-type="bibr" rid="B36-tijfs-8-4">36</xref>]. It determines the percentage of films that are accurately classified as fraudulent (True Positives) compared to all videos that are anticipated to be fraudulent (True Positives + False Positives). Our system achieved 76.70% precision. Precision shows how many of the fake videos detected are in fact fake:</p>
		  <disp-formula id="FD2-tijfs-8-4">
            <label>(2)</label>
            <mml:math id="mm2" display="block">
              <mml:semantics>
                <mml:mrow>
                  <mml:mi>P</mml:mi>
                  <mml:mi>r</mml:mi>
                  <mml:mi>e</mml:mi>
                  <mml:mi>c</mml:mi>
                  <mml:mi>i</mml:mi>
                  <mml:mi>s</mml:mi>
                  <mml:mi>i</mml:mi>
                  <mml:mi>o</mml:mi>
                  <mml:mi>n</mml:mi>
                  <mml:mo>=</mml:mo>
                  <mml:mstyle scriptlevel="0" displaystyle="true">
                    <mml:mfrac>
                      <mml:mrow>
                        <mml:mi>T</mml:mi>
                        <mml:mi>P</mml:mi>
                      </mml:mrow>
                      <mml:mrow>
                        <mml:mi>T</mml:mi>
                        <mml:mi>P</mml:mi>
                        <mml:mo>+</mml:mo>
                        <mml:mi>F</mml:mi>
                        <mml:mi>P</mml:mi>
                      </mml:mrow>
                    </mml:mfrac>
                  </mml:mstyle>
                </mml:mrow>
              </mml:semantics>
            </mml:math>
          </disp-formula>
        </sec>
        <sec id="sec2dot3dot3-tijfs-8-4">
          <title>2.3.3. Recall</title>
          <p>The model&#x2019;s recall measures its capacity to recognize every fake video in the collection [<xref ref-type="bibr" rid="B37-tijfs-8-4">37</xref>]. It determines the percentage of fake videos that are accurately identified (True Positives) relative to the total number of fake videos (True Positives + False Negatives). Our system achieved 89.15% recall. Recall shows how many actual fake videos were correctly identified:</p>
		  <disp-formula id="FD3-tijfs-8-4">
            <label>(3)</label>
            <mml:math id="mm3" display="block">
              <mml:semantics>
                <mml:mrow>
                  <mml:mi>R</mml:mi>
                  <mml:mi>e</mml:mi>
                  <mml:mi>c</mml:mi>
                  <mml:mi>a</mml:mi>
                  <mml:mi>l</mml:mi>
                  <mml:mi>l</mml:mi>
                  <mml:mo>=</mml:mo>
                  <mml:mstyle scriptlevel="0" displaystyle="true">
                    <mml:mfrac>
                      <mml:mrow>
                        <mml:mi>T</mml:mi>
                        <mml:mi>P</mml:mi>
                      </mml:mrow>
                      <mml:mrow>
                        <mml:mi>T</mml:mi>
                        <mml:mi>P</mml:mi>
                        <mml:mo>+</mml:mo>
                        <mml:mi>F</mml:mi>
                        <mml:mi>N</mml:mi>
                      </mml:mrow>
                    </mml:mfrac>
                  </mml:mstyle>
                </mml:mrow>
              </mml:semantics>
            </mml:math>
          </disp-formula>
        </sec>
        <sec id="sec2dot3dot4-tijfs-8-4">
          <title>2.3.4. F1-Score</title>
          <p>The F1-score is the harmonic mean of recall and precision [<xref ref-type="bibr" rid="B38-tijfs-8-4">38</xref>]. It offers a balance between the two metrics, which is particularly helpful in situations when the distribution of classes is not uniform. Our system achieved an F1-score of 82.46%. The F1-score is a balance between precision and recall, calculated as</p>
		  <disp-formula id="FD4-tijfs-8-4">
            <label>(4)</label>
            <mml:math id="mm4" display="block">
              <mml:semantics>
                <mml:mrow>
                  <mml:mi>F</mml:mi>
                  <mml:mn>1</mml:mn>
                  <mml:mo>=</mml:mo>
                  <mml:mn>2</mml:mn>
                  <mml:mo>&#xD7;</mml:mo>
                  <mml:mstyle scriptlevel="0" displaystyle="true">
                    <mml:mfrac>
                      <mml:mrow>
                        <mml:mi>P</mml:mi>
                        <mml:mi>r</mml:mi>
                        <mml:mi>e</mml:mi>
                        <mml:mi>c</mml:mi>
                        <mml:mi>i</mml:mi>
                        <mml:mi>s</mml:mi>
                        <mml:mi>i</mml:mi>
                        <mml:mi>o</mml:mi>
                        <mml:mi>n</mml:mi>
                        <mml:mo>&#xD7;</mml:mo>
                        <mml:mi>R</mml:mi>
                        <mml:mi>e</mml:mi>
                        <mml:mi>c</mml:mi>
                        <mml:mi>a</mml:mi>
                        <mml:mi>l</mml:mi>
                        <mml:mi>l</mml:mi>
                      </mml:mrow>
                      <mml:mrow>
                        <mml:mi>P</mml:mi>
                        <mml:mi>r</mml:mi>
                        <mml:mi>e</mml:mi>
                        <mml:mi>c</mml:mi>
                        <mml:mi>i</mml:mi>
                        <mml:mi>s</mml:mi>
                        <mml:mi>i</mml:mi>
                        <mml:mi>o</mml:mi>
                        <mml:mi>n</mml:mi>
                        <mml:mo>+</mml:mo>
                        <mml:mi>R</mml:mi>
                        <mml:mi>e</mml:mi>
                        <mml:mi>c</mml:mi>
                        <mml:mi>a</mml:mi>
                        <mml:mi>l</mml:mi>
                        <mml:mi>l</mml:mi>
                      </mml:mrow>
                    </mml:mfrac>
                  </mml:mstyle>
                </mml:mrow>
              </mml:semantics>
            </mml:math>
          </disp-formula>
          <p>The final system allows the user to upload the video through a web interface, which is developed in Streamlit (<uri>https://streamlit.io/</uri>), Ver 1.59.1. The uploaded file is then processed in real time, extracting frames from the videos by cropping, resizing, and normalizing them. Streamlit, in which the web application is developed, allows users to upload videos and view the results after processing. The backend handles the preprocessing, the loading of the model, and real-time prediction for the user.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec3-tijfs-8-4" sec-type="results">
      <title>3. Results</title>
      <p>The Google Colab environment&#x2019;s TensorFlow and Keras frameworks were used for the training procedure. A batch size of 32 samples was used to train the system across 30 epochs. Model weights were updated during training using the Adam optimizer, which was chosen due to its effectiveness while working with sparse gradients and large datasets. During training, a constant learning rate of 0.001 was used. The binary cross-entropy loss function was employed because it works well for two-class classification issues. In general, the training setup was chosen to minimize overfitting and provide strong convergence.</p>
      <p><xref ref-type="fig" rid="tijfs-8-4-f003">Figure 3</xref> demonstrates the model&#x2019;s ability to discriminate between real and false frames, as shown by the high number of true positives and true negatives. There were a few instances of false positives, in which genuine frames were mistakenly identified as fraudulent.</p>
	  <fig id="tijfs-8-4-f003" position="anchor">
        <label>Figure 3</label>
        <caption>
          <p>The confusion matrix shows a real vs. fake classifier with 20,092 true negatives, 20,453 true positives, 6213 false positives, and 2488 false negatives out of 49,246 total samples. It further illustrates the classification performance of the proposed binary classification model in distinguishing between real and fake frames.</p>
        </caption>
        <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="image003.png"/>
      </fig>
      <p>The confusion matrix presented in <xref ref-type="fig" rid="tijfs-8-4-f003">Figure 3</xref> illustrates the classification performance of the proposed binary classification model in distinguishing between real and fake frames. The matrix indicates that the model correctly classified 20,092 real samples as real and 20,453 fake samples as fake, resulting in a total of 40,545 correct predictions out of 49,246 test samples. Conversely, 6213 real samples were incorrectly classified as fake, while 2488 fake samples were misclassified as real. These results correspond to an overall classification accuracy of approximately 82.33%, demonstrating that the proposed model effectively differentiates between the two classes while maintaining a relatively low misclassification rate.</p>
      <p>These results suggest that the model is more conservative in predicting the real class, thereby reducing the chances of fake samples being accepted as real. Although some genuine samples are falsely identified as fake, the overall distribution of correct predictions demonstrates that the classifier is well-balanced and reliable for practical applications. <xref ref-type="fig" rid="tijfs-8-4-f003">Figure 3</xref> confirms the robustness of the proposed model and highlights its effectiveness in accurately recognizing both classes while providing opportunities for further optimization to reduce false classifications, particularly for the real class.</p>
      <p>The following <xref ref-type="fig" rid="tijfs-8-4-f004">Figure 4</xref> shows the evaluation metrics of the deepfake detector. These metrics show in detail the classification ability of the model on the test set and allow us to determine how well the system can distinguish between real and fake videos. A high precision value indicates that most of the videos classified as fake are fake, reducing the false positive cases, while a high recall value indicates that many of the fake videos are properly identified by the model, therefore reducing the cases of false negatives. F1-score, as the harmonic mean of precision and recall, reveals the similarity in the results of those two measures, as well as the reliability of the grouping model in terms of the two classes. Lastly, there is the overall accuracy of the model with reference to the accurate classification of samples on the test set.</p>
	  <fig id="tijfs-8-4-f004" position="anchor">
        <label>Figure 4</label>
        <caption>
          <p>The model achieves an accuracy of 82%, precision of 77%, recall of 89%, and an F1 score of 82%, indicating strong overall performance with a particular strength in correctly identifying positive cases (high recall) relative to precision.</p>
        </caption>
        <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="image004.png"/>
      </fig>
      <p>The following <xref ref-type="fig" rid="tijfs-8-4-f005">Figure 5</xref> displays the model&#x2019;s training and validation losses over 30 epochs. As can be observed, the training loss gradually drops, suggesting that the model is successfully picking up on trends in the data. The model appears to be generalizing well with new data, as evidenced by the validation loss&#x2019;s comparable trend and its close alignment with the training curve. One indication that there is little overfitting is the lack of significant gaps or oscillations between the two curves.</p>
	  <fig id="tijfs-8-4-f005" position="anchor">
        <label>Figure 5</label>
        <caption>
          <p>Training and validation loss over epochs.</p>
        </caption>
        <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="image005.png"/>
      </fig>
      <p>Similarly, <xref ref-type="fig" rid="tijfs-8-4-f006">Figure 6</xref> below further illustrates the training and validation accuracy of the proposed ResNet-50 model over 27 training epochs. The graph demonstrates a steady improvement in classification performance throughout the training process, indicating that the model effectively learns meaningful feature representations from the training data. The training accuracy increases consistently from approximately 53% in the first epoch to 80% in the final epoch, reflecting gradual convergence and stable optimization. Similarly, the validation accuracy improves from approximately 56% to 83%, showing that the model generalizes well to previously unseen data.</p>
      <p>It should be noted that the validation accuracy remains slightly higher than the training accuracy during most training epochs. This behavior suggests that the model benefits from effective regularization techniques, such as data augmentation, dropout, or batch normalization, which help prevent overfitting and improve generalization performance. Moreover, the relatively small gap between the two curves indicates that the model has achieved a good balance between learning the training data and maintaining strong predictive performance on the validation set.</p>
	  <fig id="tijfs-8-4-f006" position="anchor">
        <label>Figure 6</label>
        <caption>
          <p>Training and validation accuracy over epochs. Both training and validation loss decreased steadily over 27 epochs, from around 0.71 and 0.66 to approximately 0.40 and 0.35, respectively, indicating consistent learning without signs of overfitting, as the validation loss remains close to and even below the training loss throughout.</p>
        </caption>
        <graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="image006.png"/>
      </fig>
      <p>A steady rising trend is seen in both curves, suggesting that performance has improved over time. Training that is balanced and successful is indicated by both curves increasing steadily and without divergence.</p>
      <sec id="sec3dot1-tijfs-8-4">
        <title>3.1. Model Performance Across Selected Epochs</title>
        <p>The model training history at certain epochs, the first, fifth, tenth, fifteenth, twenty-fifth, and twenty-seventh epochs, is displayed in <xref ref-type="table" rid="tijfs-8-4-t003">Table 3</xref>.</p>
		<table-wrap id="tijfs-8-4-t003" position="anchor">
        <object-id pub-id-type="pii">tijfs-8-4-t003_Table 3</object-id>
        <label>Table 3</label>
        <caption>
          <p>Model training history at selected epochs.</p>
        </caption>
        <table>
          <thead>
            <tr>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Epoch</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Accuracy</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Val Accuracy</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Loss</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Val Loss</th>
            </tr>
          </thead>
          <tbody>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>1</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.53</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.56</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.71</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.66</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>5</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.58</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.63</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.66</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.63</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>10</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.67</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.71</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.57</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.55</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>15</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.73</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.76</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.52</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.48</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>20</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.76</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.80</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.46</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.41</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>25</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.79</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.82</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.42</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.36</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>27</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.82</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.83</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.40</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.35</td>
            </tr>
          </tbody>
        </table>
      </table-wrap>
        <p><xref ref-type="table" rid="tijfs-8-4-t003">Table 3</xref> provides the results for training loss, validation loss, training accuracy, and validation accuracy. On both the training and validation datasets, the model&#x2019;s accuracy steadily increased as training went on, while the loss values continuously dropped. These findings show that the model was learning efficiently and performing well in generalization on data that had not yet been seen.</p>
        <sec id="sec3dot1dot1-tijfs-8-4">
          <title>3.1.1. Model Performance Comparison</title>
          <p>The prior study shown in <xref ref-type="table" rid="tijfs-8-4-t004">Table 4</xref> relies on the ResNet-50 model, but uses it with different data and training strategies. In that study, the focus lies on image-level detection, followed by using preprocessing methods, i.e., Local Binary Pattern (LBP) and Gaussian filtering, to enhance feature extraction and the corresponding performance. The second study focuses on video-level classification, where a second layer of binary classification is added, and a binary cross-entropy loss is initialized to learn from a larger amount of classification data. The comparison generally reveals that the size of the dataset and preprocessing and training approaches are also important in identifying the authenticity of deepfake detection models. The model comparison is given in <xref ref-type="table" rid="tijfs-8-4-t004">Table 4</xref>.</p>
		  <table-wrap id="tijfs-8-4-t004" position="anchor">
        <object-id pub-id-type="pii">tijfs-8-4-t004_Table 4</object-id>
        <label>Table 4</label>
        <caption>
          <p>Model comparison.</p>
        </caption>
        <table>
          <thead>
            <tr>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Author</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Dataset</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Methodology</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Model<break/>Used</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Simulation Environment</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Conclusion</th>
            </tr>
          </thead>
          <tbody>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin"><bold>Arini et al. [<xref ref-type="bibr" rid="B39-tijfs-8-4">39</xref>]</bold></td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Celeb-<break/>DF (2000 images:<break/>1000 real/1000 deepfake), JPG<break/>format, 224 &#xD7; 224 Pixels</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Gaussian Filter + LBP</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">ResNet-50</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">20 epochs, batch size 100, Adam<break/>optimizer, cross-entropy loss</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Performs well on a small dataset. LBP and filtering enhance results</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>Proposed</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">Celeb-DF<break/>which has 408 real videos and 795 fake<break/>videos</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">A binary classification layer is added on top + binary<break/>cross-entropy</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">ResNet-50</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">30 epochs, batch size is 32, Adam<break/>optimizer</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">perform<break/>Well, on the large dataset, and enhance the accuracy.</td>
            </tr>
          </tbody>
        </table>
      </table-wrap>
        </sec>
        <sec id="sec3dot1dot2-tijfs-8-4">
          <title>3.1.2. Performance Metrics Comparison</title>
          <p>As can be seen in <xref ref-type="table" rid="tijfs-8-4-t001">Table 1</xref>, the proposed model based on ResNet-50 performed well and showed an accuracy of 0.82, which indicates the overall correctness of the prediction. These metrics of 0.76 and 0.89 indicate that the majority of detected deepfakes were reported correctly, and the model was very efficient in identifying most of the real deepfake instances. In addition, the F1-score of 0.82, which is balanced with precision and recall, is a further indication of the stability of the model because it ensures both high accuracy and coherence in various cases.</p>
          <p>These findings indicate that the ResNet-50 network shows an appropriate balance and offers an effective deepfake-detection option overall. The following is the performance metrics comparison, which is given in <xref ref-type="table" rid="tijfs-8-4-t005">Table 5</xref>.</p>
		  <table-wrap id="tijfs-8-4-t005" position="anchor">
        <object-id pub-id-type="pii">tijfs-8-4-t005_Table 5</object-id>
        <label>Table 5</label>
        <caption>
          <p>Performance metrics of the ResNet-50 model.</p>
        </caption>
        <table>
          <thead>
            <tr>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin"> </th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Model</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Accuracy</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Precision</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">Recall</th>
              <th align="left" valign="middle" style="border-top:solid thin;border-bottom:solid thin">F1-Score</th>
            </tr>
          </thead>
          <tbody>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>Existing</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">ResNet-50</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.75</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.80</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.68</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.77</td>
            </tr>
            <tr>
              <td align="left" valign="middle" style="border-bottom:solid thin">
                <bold>Proposed</bold>
              </td>
              <td align="left" valign="middle" style="border-bottom:solid thin">ResNet-50</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.82.33</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.76.70</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.89.15</td>
              <td align="left" valign="middle" style="border-bottom:solid thin">0.82.46</td>
            </tr>
          </tbody>
        </table>
      </table-wrap>
        </sec>
      </sec>
    </sec>
    <sec id="sec4-tijfs-8-4" sec-type="discussion">
      <title>4. Discussion</title>
      <p>The proposed model achieved 82.33% accuracy, 76.70% precision, 89.15% recall, and an 82.46% F1-score. These evaluation results show that the model performs well in distinguishing between real and fake videos. The training and validation curves represent stable learning behavior without signs of overfitting, while the confusion matrix confirms the high number of correct predictions. We propose the use of this system for real-world deployments in media forensics and content moderation platforms. The outcomes and implementation specifics of the deepfake detection system have been shown, and we started by explaining the experiment configuration, namely the instruments, libraries, and hardware used for performing the DL model training and testing process.</p>
      <p><xref ref-type="table" rid="tijfs-8-4-t005">Table 5</xref> presents a comparative performance analysis of the existing and proposed ResNet-50 models using four standard evaluation metrics: accuracy, precision, recall, and F1-score. The results demonstrate that the proposed model consistently outperforms the baseline ResNet-50 across all major performance indicators. Specifically, the existing ResNet-50 achieved an accuracy of 75.00%, whereas the proposed approach improved the classification accuracy to 82.33%, representing a significant enhancement in the model&#x2019;s overall predictive capability.</p>
      <p>Furthermore, the proposed model achieved a precision of 76.70%, a recall of 89.15%, and an F1-score of 82.46%, compared with the existing ResNet-50, which obtained 80.00%, 68.00%, and 77.00%, respectively. Although the proposed model exhibits a slightly lower precision than the baseline, it demonstrates a substantial improvement in recall, indicating a much stronger ability to correctly identify positive instances while significantly reducing false negatives. The higher F1-score confirms that the proposed approach provides a better balance between precision and recall, leading to superior overall classification performance. These results validate the effectiveness of the proposed enhancements to the ResNet-50 architecture and demonstrate its suitability for accurate and reliable classification in the target application.</p>
      <p>The surroundings were properly organized to be acceptable, so that the system could be efficiently trained and tested precisely. The evaluation metrics were accuracy, precision, recall, and F1-score. They assisted in assessing the effectiveness of the model in identifying deepfake videos versus real videos. A confusion matrix was also used to provide a clear insight into the correct or wrong number of videos that were identified as real or fake. It assisted in determining the strengths and weaknesses of the model, in particular, where the model went wrong. The training and validation accuracy were also plotted, demonstrating the level of model performance during training and on validation data. These outcomes proved that the model was not overfitting.</p>
    </sec>
    <sec id="sec5-tijfs-8-4" sec-type="conclusions">
      <title>5. Conclusions</title>
      <p>This study introduces a deepfake video detection system based on DL, namely a CNN model based on ResNet50. The utilization of deepfakes represents a significant problem because it leads to false information, reputation destruction, a potential impact on national security. Our system includes preprocessing of videos by the extraction of certain important frames, identifying and scaling faces, and converting them to a suitable format. The ResNet50 model used has a particularly strong architecture in terms of its image recognition abilities, and the model&#x2019;s performance, indicated by the accuracy, precision, recall, and F1-score, was reasonably good. Although there are limitations to this study, such as the large amounts of data and high-performance equipment required, this study shows that DL is potentially useful when combating fake media and misinformation.</p>
      <p>However, within the scope of this work on detecting deepfake videos through the classification and visual analysis of images, there are many domains in which improvements can be achieved. Another potential direction of study is related to introducing audio analysis, because many deepfakes also exploit speech editing by using AI-generated voice models. Future research may focus on speech patterns, tone, frequency, and lip sync to analyze videos using DL and signal processing, aiming for higher accuracy in detecting elements such as inconsistencies between audio and audiovisual data. Another critical area is the improvement of the user interface.</p>
    </sec>
  </body>
  <back>
    <notes>
      <title>Author Contributions</title>
      <p>Conceptualization, I.A.; methodology, M.A.K. and E.U.; software, M.A.K. and E.U.; validation, I.A., M.A.K. and E.U.; formal analysis, I.A.; investigation, M.A.K. and E.U.; resources, I.A.; data curation, M.A.K. and E.U.; writing&#x2014;original draft preparation, M.A.K. and E.U.; writing&#x2014;review and editing, I.A.; visualization, M.A.K. and E.U.; supervision, I.A.; project administration, I.A.; funding acquisition, N/A. All authors have read and agreed to the published version of the manuscript.</p>
    </notes>
	<notes>
      <title>Funding</title>
	  <p>This study received no external funding from any source.</p>
    </notes>
    <notes>
      <title>Data Availability Statement</title>
      <p>The source code, implementation, and user interface (UI) developed for this study are publicly available to facilitate reproducibility, validation, and further research. Interested researchers can access the frontend and backend repositories through the following GitHub links: Frontend (User Interface): <uri>https://github.com/ehaab212/fake-video-detection-frontend</uri> (accessed on 2 August 2026); Backend (Model and API): <uri>https://github.com/ehaab212/fake-video-detection-backend</uri> (accessed on 2 August 2026).</p>
    </notes>
    <notes notes-type="COI-statement">
      <title>Conflicts of Interest</title>
      <p>The authors claim no conflicts of interest.</p>
    </notes>
    <ref-list>
      <title>References</title>
      <ref id="B1-tijfs-8-4">
        <label>1.</label>
        <element-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Chadha</surname>
              <given-names>A.</given-names>
            </name>
            <name>
              <surname>Kumar</surname>
              <given-names>V.</given-names>
            </name>
            <name>
              <surname>Kashyap</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Gupta</surname>
              <given-names>M.</given-names>
            </name>
          </person-group>
          <article-title>Deepfake: An Overview</article-title>
          <source>Proceedings of Second International Conference on Computing, Communications, and Cyber-Security. Lecture Notes in Networks and Systems</source>
          <person-group person-group-type="editor">
            <name>
              <surname>Singh</surname>
              <given-names>P.K.</given-names>
            </name>
            <name>
              <surname>Wierzcho&#x144;</surname>
              <given-names>S.T.</given-names>
            </name>
            <name>
              <surname>Tanwar</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Ganzha</surname>
              <given-names>M.</given-names>
            </name>
            <name>
              <surname>Rodrigues</surname>
              <given-names>J.J.P.C.</given-names>
            </name>
          </person-group>
          <publisher-name>Springer</publisher-name>
          <publisher-loc>Singapore</publisher-loc>
          <year>2021</year>
          <volume>Volume 203</volume>
          <pub-id pub-id-type="doi">10.1007/978-981-16-0733-2_39</pub-id>
        </element-citation>
      </ref>
      <ref id="B2-tijfs-8-4">
        <label>2.</label>
        <element-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Gaur</surname>
              <given-names>L.</given-names>
            </name>
            <name>
              <surname>Arora</surname>
              <given-names>G.K.</given-names>
            </name>
            <name>
              <surname>Jhanjhi</surname>
              <given-names>N.Z.</given-names>
            </name>
          </person-group>
          <article-title>Deep learning techniques for creation of deepfakes</article-title>
          <source>DeepFakes</source>
          <publisher-name>CRC Press</publisher-name>
          <publisher-loc>Boca Raton, FL, USA</publisher-loc>
          <year>2022</year>
          <fpage>23</fpage>
          <lpage>34</lpage>
          <pub-id pub-id-type="doi">10.1201/9781003231493-3</pub-id>
        </element-citation>
      </ref>
      <ref id="B3-tijfs-8-4">
        <label>3.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Asim</surname>
              <given-names>M.</given-names>
            </name>
            <name>
              <surname>Waqar</surname>
              <given-names>M.</given-names>
            </name>
            <name>
              <surname>Alam</surname>
              <given-names>I.</given-names>
            </name>
          </person-group>
          <article-title>A Comparative Analysis of Deep Learning Methods for Slang Detection in Twitter Data</article-title>
          <source>Spectr. Eng. Sci.</source>
          <year>2025</year>
          <volume>3</volume>
          <issue>12</issue>
          <fpage>254</fpage>
          <lpage>270</lpage>
          <pub-id pub-id-type="doi">10.5281/zenodo.17906419</pub-id>
        </element-citation>
      </ref>
      <ref id="B4-tijfs-8-4">
        <label>4.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Yessimova</surname>
              <given-names>M.</given-names>
            </name>
            <name>
              <surname>Shevyakova</surname>
              <given-names>T.</given-names>
            </name>
          </person-group>
          <article-title>Deep Fakes in the Digital Media Age: Opportunities and Threats</article-title>
          <source>Her. J. &#x17D;urnalistika Seri&#xE2;sy</source>
          <year>2024</year>
          <volume>73</volume>
          <issue>3</issue>
          <fpage>44</fpage>
          <lpage>54</lpage>
          <pub-id pub-id-type="doi">10.26577/HJ.2024.v73.i3.4</pub-id>
        </element-citation>
      </ref>
      <ref id="B5-tijfs-8-4">
        <label>5.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Majerczak</surname>
              <given-names>P.</given-names>
            </name>
            <name>
              <surname>Strzelecki</surname>
              <given-names>A.</given-names>
            </name>
          </person-group>
          <article-title>Trust, media credibility, social ties, and the intention to share towards information verification in an age of fake news</article-title>
          <source>Behav. Sci.</source>
          <year>2022</year>
          <volume>12</volume>
          <issue>2</issue>
          <elocation-id>51</elocation-id>
          <pub-id pub-id-type="doi">10.3390/bs12020051</pub-id>
        </element-citation>
      </ref>
      <ref id="B6-tijfs-8-4">
        <label>6.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Mirsky</surname>
              <given-names>Y.</given-names>
            </name>
            <name>
              <surname>Lee</surname>
              <given-names>W.</given-names>
            </name>
          </person-group>
          <article-title>The creation and detection of deepfakes: A survey</article-title>
          <source>ACM Comput. Surv. (CSUR)</source>
          <year>2021</year>
          <volume>54</volume>
          <issue>1</issue>
          <fpage>1</fpage>
          <lpage>41</lpage>
          <pub-id pub-id-type="doi">10.1145/3425780</pub-id>
        </element-citation>
      </ref>
      <ref id="B7-tijfs-8-4">
        <label>7.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Fatima</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Ali</surname>
              <given-names>Z.</given-names>
            </name>
            <name>
              <surname>Alam</surname>
              <given-names>I.</given-names>
            </name>
            <name>
              <surname>Muhammad</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Jan</surname>
              <given-names>G.</given-names>
            </name>
          </person-group>
          <article-title>Unmasking Hate: Deep Learning-Based Detection of Violent Incitement in Social Media</article-title>
          <source>Annu. Methodol. Arch. Res. Rev.</source>
          <year>2025</year>
          <volume>3</volume>
          <issue>12</issue>
          <fpage>138</fpage>
          <lpage>160</lpage>
          <pub-id pub-id-type="doi">10.63075/k80ymr06</pub-id>
        </element-citation>
      </ref>
      <ref id="B8-tijfs-8-4">
        <label>8.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Goh</surname>
              <given-names>D.H.L.</given-names>
            </name>
          </person-group>
          <article-title>&#x201C;He looks very real&#x201D;: Media, knowledge, and search-based strategies for deepfake identification</article-title>
          <source>J. Assoc. Inf. Sci. Technol.</source>
          <year>2024</year>
          <volume>75</volume>
          <issue>6</issue>
          <fpage>643</fpage>
          <lpage>654</lpage>
          <pub-id pub-id-type="doi">10.1002/asi.24867</pub-id>
        </element-citation>
      </ref>
      <ref id="B9-tijfs-8-4">
        <label>9.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Jin</surname>
              <given-names>X.</given-names>
            </name>
            <name>
              <surname>Yi</surname>
              <given-names>K.</given-names>
            </name>
            <name>
              <surname>Xu</surname>
              <given-names>J.</given-names>
            </name>
          </person-group>
          <article-title>MoADNet: Mobile asymmetric dual-stream networks for real-time and lightweight RGB-D salient object detection</article-title>
          <source>IEEE Trans. Circuits Syst. Video Technol.</source>
          <year>2022</year>
          <volume>32</volume>
          <issue>11</issue>
          <fpage>7632</fpage>
          <lpage>7645</lpage>
          <pub-id pub-id-type="doi">10.1109/TCSVT.2022.3180274</pub-id>
        </element-citation>
      </ref>
      <ref id="B10-tijfs-8-4">
        <label>10.</label>
        <element-citation publication-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Sharma</surname>
              <given-names>V.K.</given-names>
            </name>
            <name>
              <surname>Rawat</surname>
              <given-names>S.</given-names>
            </name>
          </person-group>
          <article-title>Enhancing Deepfake Detection Through Dynamics of Facial Expressions</article-title>
          <source>Proceedings of the 2025 6th International Conference on Intelligent Communication Technologies and Virtual Mobile Networks (ICICV)</source>
          <conf-loc>Tirunelveli, India</conf-loc>
          <conf-date>17&#x2013;19 June 2025</conf-date>
          <fpage>70</fpage>
          <lpage>79</lpage>
          <pub-id pub-id-type="doi">10.1109/ICICV64824.2025.11085942</pub-id>
        </element-citation>
      </ref>
      <ref id="B11-tijfs-8-4">
        <label>11.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Mansoor</surname>
              <given-names>A.</given-names>
            </name>
            <name>
              <surname>Zaheen</surname>
              <given-names>A.</given-names>
            </name>
            <name>
              <surname>Ali</surname>
              <given-names>Z.</given-names>
            </name>
            <name>
              <surname>Idrees</surname>
              <given-names>F.</given-names>
            </name>
            <name>
              <surname>Rahim</surname>
              <given-names>M.</given-names>
            </name>
            <name>
              <surname>Jan</surname>
              <given-names>G.</given-names>
            </name>
            <name>
              <surname>Alam</surname>
              <given-names>I.</given-names>
            </name>
          </person-group>
          <article-title>Enhancing thyroid ultrasound diagnosis with a hybrid CNN and graph attention network</article-title>
          <source>Spectr. Eng. Sci.</source>
          <year>2025</year>
          <volume>3</volume>
          <fpage>95</fpage>
          <lpage>105</lpage>
          <pub-id pub-id-type="doi">10.5281/zenodo.16778001</pub-id>
        </element-citation>
      </ref>
      <ref id="B12-tijfs-8-4">
        <label>12.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Jin</surname>
              <given-names>X.</given-names>
            </name>
            <name>
              <surname>Jing</surname>
              <given-names>P.</given-names>
            </name>
            <name>
              <surname>Wu</surname>
              <given-names>J.</given-names>
            </name>
            <name>
              <surname>Xu</surname>
              <given-names>J.</given-names>
            </name>
            <name>
              <surname>Su</surname>
              <given-names>Y.</given-names>
            </name>
          </person-group>
          <article-title>Visual sentiment classification via low-rank regularization and label relaxation</article-title>
          <source>IEEE Trans. Cogn. Dev. Syst.</source>
          <year>2021</year>
          <volume>14</volume>
          <issue>4</issue>
          <fpage>1678</fpage>
          <lpage>1690</lpage>
          <pub-id pub-id-type="doi">10.1109/TCDS.2021.3135948</pub-id>
        </element-citation>
      </ref>
      <ref id="B13-tijfs-8-4">
        <label>13.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Jin</surname>
              <given-names>X.</given-names>
            </name>
            <name>
              <surname>Yu</surname>
              <given-names>W.</given-names>
            </name>
            <name>
              <surname>Chen</surname>
              <given-names>D.-W.</given-names>
            </name>
            <name>
              <surname>Shi</surname>
              <given-names>W.</given-names>
            </name>
          </person-group>
          <article-title>DFD-NAS: General deepfake detection via efficient neural architecture search</article-title>
          <source>Neurocomputing</source>
          <year>2025</year>
          <volume>619</volume>
          <fpage>129129</fpage>
          <pub-id pub-id-type="doi">10.1016/j.neucom.2024.129129</pub-id>
        </element-citation>
      </ref>
      <ref id="B14-tijfs-8-4">
        <label>14.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Alam</surname>
              <given-names>I.</given-names>
            </name>
            <name>
              <surname>Basit</surname>
              <given-names>A.</given-names>
            </name>
            <name>
              <surname>Ziar</surname>
              <given-names>R.A.</given-names>
            </name>
          </person-group>
          <article-title>Utilizing Age-Adaptive Deep Learning Approaches for Detecting Inappropriate Video Content</article-title>
          <source>Hum. Behav. Emerg. Technol.</source>
          <year>2024</year>
          <volume>2024</volume>
          <issue>1</issue>
          <fpage>7004031</fpage>
          <pub-id pub-id-type="doi">10.1155/2024/7004031</pub-id>
        </element-citation>
      </ref>
      <ref id="B15-tijfs-8-4">
        <label>15.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Jin</surname>
              <given-names>X.</given-names>
            </name>
            <name>
              <surname>Guo</surname>
              <given-names>C.</given-names>
            </name>
            <name>
              <surname>He</surname>
              <given-names>Z.</given-names>
            </name>
            <name>
              <surname>Xu</surname>
              <given-names>J.</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>Y.</given-names>
            </name>
            <name>
              <surname>Su</surname>
              <given-names>Y.</given-names>
            </name>
          </person-group>
          <article-title>FCMNet: Frequency-aware cross-modality attention networks for RGB-D salient object detection</article-title>
          <source>Neurocomputing</source>
          <year>2022</year>
          <volume>491</volume>
          <fpage>414</fpage>
          <lpage>425</lpage>
          <pub-id pub-id-type="doi">10.1016/j.neucom.2022.04.015</pub-id>
        </element-citation>
      </ref>
      <ref id="B16-tijfs-8-4">
        <label>16.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Karandikar</surname>
              <given-names>A.</given-names>
            </name>
            <name>
              <surname>Deshpande</surname>
              <given-names>V.</given-names>
            </name>
            <name>
              <surname>Singh</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Nagbhidkar</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Agrawal</surname>
              <given-names>S.</given-names>
            </name>
          </person-group>
          <article-title>Deepfake video detection using convolutional neural network</article-title>
          <source>Int. J. Adv. Trends Comput. Sci. Eng.</source>
          <year>2020</year>
          <volume>9</volume>
          <fpage>1311</fpage>
          <lpage>1315</lpage>
          <pub-id pub-id-type="doi">10.30534/ijatcse/2020/62922020</pub-id>
        </element-citation>
      </ref>
      <ref id="B17-tijfs-8-4">
        <label>17.</label>
        <element-citation publication-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Zhao</surname>
              <given-names>H.</given-names>
            </name>
            <name>
              <surname>Zhou</surname>
              <given-names>W.</given-names>
            </name>
            <name>
              <surname>Chen</surname>
              <given-names>D.</given-names>
            </name>
            <name>
              <surname>Wei</surname>
              <given-names>T.</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>W.</given-names>
            </name>
            <name>
              <surname>Yu</surname>
              <given-names>N.</given-names>
            </name>
          </person-group>
          <article-title>Multi-attentional deepfake detection</article-title>
          <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          <conf-loc>Nashville, TN, USA</conf-loc>
          <conf-date>19&#x2013;25 June 2021</conf-date>
          <fpage>2185</fpage>
          <lpage>2194</lpage>
          <pub-id pub-id-type="doi">10.1109/CVPR46437.2021.00222</pub-id>
        </element-citation>
      </ref>
      <ref id="B18-tijfs-8-4">
        <label>18.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Doke</surname>
              <given-names>Y.</given-names>
            </name>
            <name>
              <surname>Dongare</surname>
              <given-names>P.</given-names>
            </name>
            <name>
              <surname>Marathe</surname>
              <given-names>V.</given-names>
            </name>
            <name>
              <surname>Gaikwad</surname>
              <given-names>M.</given-names>
            </name>
            <name>
              <surname>Gaikwad</surname>
              <given-names>M.</given-names>
            </name>
          </person-group>
          <article-title>Deepfake video detection using deep learning</article-title>
          <source>Int. J. Res. Publ. Rev.</source>
          <year>2022</year>
          <volume>2582</volume>
          <fpage>7421</fpage>
        </element-citation>
      </ref>
      <ref id="B19-tijfs-8-4">
        <label>19.</label>
        <element-citation publication-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Mitra</surname>
              <given-names>A.</given-names>
            </name>
            <name>
              <surname>Mohanty</surname>
              <given-names>S.P.</given-names>
            </name>
            <name>
              <surname>Corcoran</surname>
              <given-names>P.</given-names>
            </name>
            <name>
              <surname>Kougianos</surname>
              <given-names>E.</given-names>
            </name>
          </person-group>
          <article-title>A novel machine learning based method for deepfake video detection in social media</article-title>
          <source>Proceedings of the 2020 IEEE International Symposium on Smart Electronic Systems (iSES) (Formerly INiS)</source>
          <conf-loc>Chennai, India</conf-loc>
          <conf-date>14&#x2013;16 December 2020</conf-date>
          <fpage>91</fpage>
          <lpage>96</lpage>
          <pub-id pub-id-type="doi">10.1109/iSES50453.2020.00031</pub-id>
        </element-citation>
      </ref>
      <ref id="B20-tijfs-8-4">
        <label>20.</label>
        <element-citation publication-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Bonettini</surname>
              <given-names>N.</given-names>
            </name>
            <name>
              <surname>Cannas</surname>
              <given-names>E.D.</given-names>
            </name>
            <name>
              <surname>Mandelli</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Bondi</surname>
              <given-names>L.</given-names>
            </name>
            <name>
              <surname>Bestagini</surname>
              <given-names>P.</given-names>
            </name>
            <name>
              <surname>Tubaro</surname>
              <given-names>S.</given-names>
            </name>
          </person-group>
          <article-title>Video face manipulation detection through ensemble of cnns</article-title>
          <source>Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR)</source>
          <conf-loc>Milan, Italy</conf-loc>
          <conf-date>10&#x2013;15 January 2021</conf-date>
          <fpage>5012</fpage>
          <lpage>5019</lpage>
          <pub-id pub-id-type="doi">10.1109/ICPR48806.2021.9412711</pub-id>
        </element-citation>
      </ref>
      <ref id="B21-tijfs-8-4">
        <label>21.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Suratkar</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Kazi</surname>
              <given-names>F.</given-names>
            </name>
          </person-group>
          <article-title>Deep fake video detection using transfer learning approach</article-title>
          <source>Arab. J. Sci. Eng.</source>
          <year>2023</year>
          <volume>48</volume>
          <issue>8</issue>
          <fpage>9727</fpage>
          <lpage>9737</lpage>
          <pub-id pub-id-type="doi">10.1007/s13369-022-07321-3</pub-id>
        </element-citation>
      </ref>
      <ref id="B22-tijfs-8-4">
        <label>22.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Patel</surname>
              <given-names>N.</given-names>
            </name>
            <name>
              <surname>Jethwa</surname>
              <given-names>N.</given-names>
            </name>
            <name>
              <surname>Mali</surname>
              <given-names>C.</given-names>
            </name>
            <name>
              <surname>Deone</surname>
              <given-names>J.</given-names>
            </name>
          </person-group>
          <article-title>Deepfake video detection using neural networks</article-title>
          <source>ITM Web Conf.</source>
          <year>2022</year>
          <volume>44</volume>
          <fpage>03024</fpage>
          <pub-id pub-id-type="doi">10.1051/itmconf/20224403024</pub-id>
        </element-citation>
      </ref>
      <ref id="B23-tijfs-8-4">
        <label>23.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Guhagarkar</surname>
              <given-names>N.</given-names>
            </name>
            <name>
              <surname>Desai</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Vaishyampayan</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Save</surname>
              <given-names>A.</given-names>
            </name>
          </person-group>
          <article-title>Deepfake Detection Techniques: A Review</article-title>
          <source>VIVA-Tech Int. J. Res. Innov. IJRI</source>
          <year>2021</year>
          <volume>1</volume>
          <issue>4</issue>
          <fpage>1</fpage>
          <lpage>10</lpage>
          <pub-id pub-id-type="doi">10.1561/116.00000024</pub-id>
        </element-citation>
      </ref>
      <ref id="B24-tijfs-8-4">
        <label>24.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Soudy</surname>
              <given-names>A.H.</given-names>
            </name>
            <name>
              <surname>Sayed</surname>
              <given-names>O.</given-names>
            </name>
            <name>
              <surname>Tag-Elser</surname>
              <given-names>H.</given-names>
            </name>
            <name>
              <surname>Ragab</surname>
              <given-names>R.</given-names>
            </name>
            <name>
              <surname>Mohsen</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Mostafa</surname>
              <given-names>T.</given-names>
            </name>
            <name>
              <surname>Abohany</surname>
              <given-names>A.A.</given-names>
            </name>
            <name>
              <surname>Slim</surname>
              <given-names>S.O.</given-names>
            </name>
          </person-group>
          <article-title>Deepfake detection using convolutional vision transformers and convolutional neural networks</article-title>
          <source>Neural Comput. Appl.</source>
          <year>2024</year>
          <volume>36</volume>
          <issue>31</issue>
          <fpage>19759</fpage>
          <lpage>19775</lpage>
          <pub-id pub-id-type="doi">10.1007/s00521-024-10181-7</pub-id>
        </element-citation>
      </ref>
      <ref id="B25-tijfs-8-4">
        <label>25.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Heidari</surname>
              <given-names>A.</given-names>
            </name>
            <name>
              <surname>Jafari Navimipour</surname>
              <given-names>N.</given-names>
            </name>
            <name>
              <surname>Dag</surname>
              <given-names>H.</given-names>
            </name>
            <name>
              <surname>Unal</surname>
              <given-names>M.</given-names>
            </name>
          </person-group>
          <article-title>Deepfake detection using deep learning methods: A systematic and comprehensive review</article-title>
          <source>Wiley Interdiscip. Rev. Data Min. Knowl. Discov.</source>
          <year>2024</year>
          <volume>14</volume>
          <issue>2</issue>
          <fpage>e1520</fpage>
          <pub-id pub-id-type="doi">10.1002/widm.1520</pub-id>
        </element-citation>
      </ref>
      <ref id="B26-tijfs-8-4">
        <label>26.</label>
        <element-citation publication-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Ju</surname>
              <given-names>Y.</given-names>
            </name>
            <name>
              <surname>Hu</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Jia</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Chen</surname>
              <given-names>G.H.</given-names>
            </name>
            <name>
              <surname>Lyu</surname>
              <given-names>S.</given-names>
            </name>
          </person-group>
          <article-title>Improving fairness in deepfake detection</article-title>
          <source>Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision</source>
          <conf-loc>Waikoloa, HI, USA</conf-loc>
          <conf-date>3&#x2013;8 January 2024</conf-date>
          <fpage>4655</fpage>
          <lpage>4665</lpage>
          <pub-id pub-id-type="doi">10.48550/arXiv.2306.16635</pub-id>
        </element-citation>
      </ref>
      <ref id="B27-tijfs-8-4">
        <label>27.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Pei</surname>
              <given-names>G.</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>J.</given-names>
            </name>
            <name>
              <surname>Hu</surname>
              <given-names>M.</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>Z.</given-names>
            </name>
            <name>
              <surname>Wang</surname>
              <given-names>C.</given-names>
            </name>
            <name>
              <surname>Wu</surname>
              <given-names>Y.</given-names>
            </name>
            <name>
              <surname>Zhai</surname>
              <given-names>G.</given-names>
            </name>
            <name>
              <surname>Yang</surname>
              <given-names>J.</given-names>
            </name>
            <name>
              <surname>Tao</surname>
              <given-names>D.</given-names>
            </name>
          </person-group>
          <article-title>Deepfake generation and detection: A benchmark and survey</article-title>
          <source>arXiv</source>
          <year>2024</year>
          <pub-id pub-id-type="arxiv">2403.17881</pub-id>
        </element-citation>
      </ref>
      <ref id="B28-tijfs-8-4">
        <label>28.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Taeb</surname>
              <given-names>M.</given-names>
            </name>
            <name>
              <surname>Chi</surname>
              <given-names>H.</given-names>
            </name>
          </person-group>
          <article-title>Comparison of deepfake detection techniques through deep learning</article-title>
          <source>J. Cybersecur. Priv.</source>
          <year>2022</year>
          <volume>2</volume>
          <issue>1</issue>
          <fpage>89</fpage>
          <lpage>106</lpage>
          <pub-id pub-id-type="doi">10.3390/jcp2010007</pub-id>
        </element-citation>
      </ref>
      <ref id="B29-tijfs-8-4">
        <label>29.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Passos</surname>
              <given-names>L.A.</given-names>
            </name>
            <name>
              <surname>Jodas</surname>
              <given-names>D.</given-names>
            </name>
            <name>
              <surname>Costa</surname>
              <given-names>K.A.</given-names>
            </name>
            <name>
              <surname>Souza J&#xFA;nior</surname>
              <given-names>L.A.</given-names>
            </name>
            <name>
              <surname>Rodrigues</surname>
              <given-names>D.</given-names>
            </name>
            <name>
              <surname>Del Ser</surname>
              <given-names>J.</given-names>
            </name>
            <name>
              <surname>Camacho</surname>
              <given-names>D.</given-names>
            </name>
            <name>
              <surname>Papa</surname>
              <given-names>J.P.</given-names>
            </name>
          </person-group>
          <article-title>A review of deep learning-based approaches for deepfake content detection</article-title>
          <source>Expert Syst.</source>
          <year>2024</year>
          <volume>41</volume>
          <issue>8</issue>
          <fpage>e13570</fpage>
          <pub-id pub-id-type="doi">10.22541/au.169735672.27713914/v1</pub-id>
        </element-citation>
      </ref>
      <ref id="B30-tijfs-8-4">
        <label>30.</label>
        <element-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Rajagukguk</surname>
              <given-names>N.</given-names>
            </name>
            <name>
              <surname>Kencana</surname>
              <given-names>I.P.E.N.</given-names>
            </name>
            <name>
              <surname>Kusuma</surname>
              <given-names>I.G.L.W.</given-names>
            </name>
          </person-group>
          <article-title>Classification of Original and Fake Images Using Deep Learning-Resnet50</article-title>
          <source>Proceedings of the First International Conference on Applied Mathematics, Statistics, and Computing (ICAMSAC 2023)</source>
          <publisher-name>Springer Nature</publisher-name>
          <publisher-loc>Berlin, Germany</publisher-loc>
          <year>2024</year>
          <fpage>51</fpage>
          <pub-id pub-id-type="doi">10.2991/978-94-6463-413-6_6</pub-id>
        </element-citation>
      </ref>
      <ref id="B31-tijfs-8-4">
        <label>31.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Orlando</surname>
              <given-names>O.</given-names>
            </name>
            <name>
              <surname>Al Rivan</surname>
              <given-names>M.E.</given-names>
            </name>
          </person-group>
          <article-title>Klasifikasi jenis kanker kulit manusia menggunakan convolution neural network</article-title>
          <source>MDP Stud. Conf.</source>
          <year>2023</year>
          <volume>2</volume>
          <issue>1</issue>
          <fpage>144</fpage>
          <lpage>150</lpage>
          <pub-id pub-id-type="doi">10.35957/mdp-sc.v2i1.4335</pub-id>
        </element-citation>
      </ref>
      <ref id="B32-tijfs-8-4">
        <label>32.</label>
        <element-citation publication-type="web">
          <person-group person-group-type="author">
            <name>
              <surname>Notshe</surname>
              <given-names>H.</given-names>
            </name>
            <name>
              <surname>Xiao</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Svoboda</surname>
              <given-names>T.</given-names>
            </name>
          </person-group>
          <article-title>Deep Learning Deepfake Detection</article-title>
          <year>2024</year>
          <comment>Available online: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://cs231n.stanford.edu/2024/papers/deep-learning-deepfake-detection.pdf" ext-link-type="uri">https://cs231n.stanford.edu/2024/papers/deep-learning-deepfake-detection.pdf</ext-link></comment>
          <date-in-citation content-type="access-date" iso-8601-date="2026-07-18">(accessed on 18 July 2026)</date-in-citation>
        </element-citation>
      </ref>
      <ref id="B33-tijfs-8-4">
        <label>33.</label>
        <element-citation publication-type="web">
          <person-group person-group-type="author">
            <name>
              <surname>Samuel</surname>
              <given-names>F.</given-names>
            </name>
          </person-group>
          <article-title>Deep Learning vs. Shallow Learning in Detecting Anomalies in Biomedical Images</article-title>
          <source>researchgate. Net</source>
          <year>2024</year>
          <comment>Available online: <ext-link xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="https://www.researchgate.net/publication/385084822_Deep_Learning_vs_Shallow_Learning_in_Detecting_Anomalies_in_Biomedical_Images" ext-link-type="uri">https://www.researchgate.net/publication/385084822_Deep_Learning_vs_Shallow_Learning_in_Detecting_Anomalies_in_Biomedical_Images</ext-link></comment>
          <date-in-citation content-type="access-date" iso-8601-date="2026-07-18">(accessed on 18 July 2026)</date-in-citation>
        </element-citation>
      </ref>
      <ref id="B34-tijfs-8-4">
        <label>34.</label>
        <element-citation publication-type="book">
          <person-group person-group-type="author">
            <name>
              <surname>Alzahrani</surname>
              <given-names>A.</given-names>
            </name>
            <name>
              <surname>Rawat</surname>
              <given-names>D.B.</given-names>
            </name>
          </person-group>
          <article-title>Enhance Deepfake Video Detection Through Optical Flow Algorithms-Based CNN</article-title>
          <source>International Conference on Human-Computer Interaction</source>
          <publisher-name>Springer</publisher-name>
          <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>
          <year>2024</year>
          <fpage>14</fpage>
          <lpage>22</lpage>
          <pub-id pub-id-type="doi">10.1007/978-3-031-62110-9_2</pub-id>
        </element-citation>
      </ref>
      <ref id="B35-tijfs-8-4">
        <label>35.</label>
        <element-citation publication-type="journal">
          <person-group person-group-type="author">
            <name>
              <surname>Bankar</surname>
              <given-names>V.</given-names>
            </name>
            <name>
              <surname>Pawar</surname>
              <given-names>D.R.</given-names>
            </name>
            <name>
              <surname>Yannawar</surname>
              <given-names>P.L.</given-names>
            </name>
          </person-group>
          <article-title>Hybrid ResNeXt and LSTM Model for Enhanced Deepfake Detection on the FaceForensics++ Dataset</article-title>
          <source>Technology</source>
          <year>2025</year>
          <volume>13</volume>
          <fpage>14</fpage>
          <pub-id pub-id-type="doi">10.17577/IJERTV14IS120061</pub-id>
        </element-citation>
      </ref>
      <ref id="B36-tijfs-8-4">
        <label>36.</label>
        <element-citation publication-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Li</surname>
              <given-names>Y.</given-names>
            </name>
            <name>
              <surname>Yang</surname>
              <given-names>X.</given-names>
            </name>
            <name>
              <surname>Sun</surname>
              <given-names>P.</given-names>
            </name>
            <name>
              <surname>Qi</surname>
              <given-names>H.</given-names>
            </name>
            <name>
              <surname>Lyu</surname>
              <given-names>S.</given-names>
            </name>
          </person-group>
          <article-title>Celeb-df: A large-scale challenging dataset for deepfake forensics</article-title>
          <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          <conf-loc>Seattle, WA, USA</conf-loc>
          <conf-date>13&#x2013;19 June 2020</conf-date>
          <fpage>3207</fpage>
          <lpage>3216</lpage>
          <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00327</pub-id>
        </element-citation>
      </ref>
      <ref id="B37-tijfs-8-4">
        <label>37.</label>
        <element-citation publication-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>He</surname>
              <given-names>K.</given-names>
            </name>
            <name>
              <surname>Zhang</surname>
              <given-names>X.</given-names>
            </name>
            <name>
              <surname>Ren</surname>
              <given-names>S.</given-names>
            </name>
            <name>
              <surname>Sun</surname>
              <given-names>J.</given-names>
            </name>
          </person-group>
          <article-title>Deep residual learning for image recognition</article-title>
          <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          <conf-loc>Las Vegas, NV, USA</conf-loc>
          <conf-date>27&#x2013;30 June 2016</conf-date>
          <fpage>770</fpage>
          <lpage>778</lpage>
          <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id>
        </element-citation>
      </ref>
      <ref id="B38-tijfs-8-4">
        <label>38.</label>
        <element-citation publication-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Yacouby</surname>
              <given-names>R.</given-names>
            </name>
            <name>
              <surname>Axman</surname>
              <given-names>D.</given-names>
            </name>
          </person-group>
          <article-title>Probabilistic extension of precision, recall, and f1 score for more thorough evaluation of classification models</article-title>
          <source>Proceedings of the First Workshop on Evaluation and Comparison of NLP Systems</source>
          <conf-loc>Online</conf-loc>
          <conf-date>20 November 2020</conf-date>
          <fpage>79</fpage>
          <lpage>91</lpage>
          <pub-id pub-id-type="doi">10.18653/v1/2020.eval4nlp-1.9</pub-id>
        </element-citation>
      </ref>
      <ref id="B39-tijfs-8-4">
        <label>39.</label>
        <element-citation publication-type="confproc">
          <person-group person-group-type="author">
            <name>
              <surname>Arini</surname>
              <given-names>A.</given-names>
            </name>
            <name>
              <surname>Bahaweres</surname>
              <given-names>R.B.</given-names>
            </name>
            <name>
              <surname>Al Haq</surname>
              <given-names>J.</given-names>
            </name>
          </person-group>
          <article-title>Quick classification of xception and resnet-50 models on deepfake video using local binary pattern</article-title>
          <source>Proceedings of the 2021 International Seminar on Machine Learning, Optimization, and Data Science (ISMODE)</source>
          <conf-loc>Jakarta, Indonesia</conf-loc>
          <conf-date>29&#x2013;30 January 2022</conf-date>
          <fpage>254</fpage>
          <lpage>259</lpage>
          <pub-id pub-id-type="doi">10.1109/ISMODE53584.2022.9742852</pub-id>
        </element-citation>
      </ref>
    </ref-list>
  </back>
</article>
