10 - read paper and provide feedback/review
DEEP-FAKE DETECTION USING CNN ENSEMBLE OF
LEARNERS WITH TRANSFER LEARNING
TECHNIQUES
ABSTRACT
Today, Computer science and its applications are growing rapidly with recent advancement of
the technology. Particularly this deep learning algorithms, image processing, fake media creations
have gained lot of recognition and creating threat to the people nowadays. Deep Fake Creation is one
amongst the large threats to the authenticity and confidential of online information. These Deep-fake will
be used for malicious intents like phishing scams and fake news. So, building this Detection system
may well be solution for fraud detection for prevention of widespread attention of fake news,
videos, footage and pictures.
The proposal of this project is to examine these media files for AI-generated media alterations in
Faces and develop new, more robust, and sophisticated approaches to deal with the more difficult
deep fakes which gets challenging everyday. The proposed architecture uses Ensemble of CNNs
containing different base learners that analyses individual face feature like mouth,eye,facial landmarks
that uses pre-trained models like ResNet, Inception-ResNet-v2 ,XceptionNet ,Efficient-Net and final
classifier predicts the best result from the learners .Experiments were conducted on data-sets that
contain both fake and real videos and also with custom built data-set of World Politicians and
Celebrities.Performance of Ensemble method has promising results when compared to other methods.
KEYWORDS
Deep learning, deep fake, Artificial intelligence, Deep-fake detection, transfer learning,
CNN,Inception-ResNet-v2.,ResNet,LSTM
1. INTRODUCTION The term "Deep-fake" refers to a technology that uses Deep Learning to create fake videos
in which one person's face is swapped/morphed onto the face of another. Deep-fakes are the
outcome of advanced machine learning and AI techniques used to manipulate or generate visual
and audio content with a high potential to deceive while imitating content. Deep learning
is primarily responsible for the creation of deep-fakes, which entails training generative
neural network designs such as auto encoders or generative adversarial networks (GANs).
GANs have the ability to create ultra-realistic fake images and videos. Various powerful
machine learning and Artificial intelligence algorithms capable of making highly realistic
Deep Fake movies are being used with malicious purpose to deceive people, harass, blackmail
women, and cause political instability by propagating false news and hostile propaganda, which
could lead to social, political, and economic instability and outbursts with terrible
consequences. Therefore, Deep fake videos pose a severe threat to personal safety, affecting not
just public images but also regular people, as well as national security, demanding the
development of automated methods to detect deep fake movies.. There are numerous public
implementations and usage of Deep fake out there, most notably Fake-App and the face swap
GitHub. Some of the Deep Fake tools are Face Forensics++, Face swap-GAN, DeepFaceLab,
Dfaker, etc. Sentinel is a start-up company
working on fake media content.\\Deep Fake Detection system has lot of applications in Forensics
for investigation of frauds, scams and help in prevention of threats, bullying, blackmails and
prevention of crime. It also helps in prevention of widespread attention of Fake news, videos,
footage and images. It reduces social abuse, harassment, promotes privacy and security of global leaders. So, finding the reality in digital domain so has become progressively essential. In fact,
when it comes to detecting deep-fake is even more challenging.\\Lately in 2019, Facebook Inc.
teaming up with Microsoft has been released the Deep-fake Detection Challenge to catalyze for
more improvement and development with better accuracy in detecting and preventing
manipulated data or deep-fakes in Kaggle and used YouTube data-set, yet remarked the challenge
is unsolved. But recent advancements in 2020, Deep Fake Detection has gained lot of popularity and lot of research have been going and quite become trending these days.Sentinel is a start-up
company working on fake media content.\\This Graph taken from https://app.dimensions.ai/
shows that at the end of August 2020 it displays that the number of papers for deep-fakes has been
increased exponentially in recent years
2. PREVIOUS WORK
2.1. DEEP LEARNING FOR DEEP FAKES CREATION AND DETECTION: A SURVEY
This research paper provides a complete survey for Deep Fake Creation and Detection which
tells about recent works, algorithms used, data-set used performance and application by various researchers. Deep Fake Detection methods are still in the initial stage and there is more
upcoming technologies in the future to detect deep-fakes. Then to summarize the table of
detection methods is displayed which uses different methods like CNN LSTM, VGG19, ResNet50, MesoNet, KNN, SVM, CFFN and many more.
2.2. DEEP FAKE DETECTION IN MEDIA FILES - AUDIOS, IMAGES AND VIDEOS
This study proposes a hybrid strategy that combines image processing and the CNN network.
There are four sections to this methodology. The first is the Data Preparation step, which
converts input samples to visual samples. The second step is the Data Enhancement section, which removes the image’s noise component. Following that is the CNN model, which is used
for training and testing to build the detection model, and finally the detection portion, which
does real-time detection against various media files from multiple platforms. Data-set: Deep-
fake Timit, Forensics++, VidTimit data-set
This paper presents a deep-fake image creation by using GAN network and there we have two networks one is generator and another one is discriminator. Here generator part of the GAN
creates fake videos and then the discriminator detects it.
2.3VIDEO DETECTION USING RECURRENT NEURAL NETWORKS (RNN)
In this research paper, they have introduced a new technique to automatically detect manipulated
videos. For extracting frame-level features they have used a CNN method. These features which
extracted from CNN method are passed to RNN that helps for classifying the given image/video has been manipulated or not. For dimensionality reduction autoencoder has been applied, LSTM
(Long short-term memory) are the one of the particular sub-types of RNN method, LSTMs are
used for temporal sequence analysis. And finally produce and probability likelihood estimation of the sequence being either a manipulated video or a real video.
Proposed Methodology: Methodologies which author used in this paper is Recurrent Neural
Network (RNN) for detect ing deepfake and for extracting frame-level features from every frame
in the video they have used CNN model. The output/features of CNN model used to train a RNN
(Recurrent Neural Network) that will learn to predict if a video/image manipulated or not. Autoencoders have been applied for dimensionality reduction and compactly represent ing the
videos/images. for temporal sequence analysis LSTM method was used.
2.4. DEEPFAKE DETECTION USING TRANSFER-LEARNING BASED CNN ARCHITECTURE
Here the authors introduce a way to detect a manipulated videos by employing a Convolutional
neural network (CNN) and using the Transferable Learning method. The technique used during this paper that one implements a method or architecture includes CNN for extracting features
from each frame of the video for training a model (binary classifier) that will learn efficiently to
differentiate the important and pretend videos. the strategy is evaluated against an extensively more set of manipulated videos those are collected from many different datasets. for generating
Deep Fakes Autoencoder method was used. Nowadays creating or modifying models for every
specific application that may be extremely difficult in terms of cost and also manpower. during this reasonably situation Transfer Learning would be helpful it’ll offer a promising solution to
such problem. the main intension of using this transferable learning method is to transfer the
knowledge which is gained from one source to other source so the understanding will be more
accurate and efficient. By using this application, we are able to build a comparatively good model for deep fake detection.
2.5. DEEPFAKE DETECTION USING CAPSULE NETWORKS WITH LONG SHORT TERM MEMORY
NETWORKS
Dataset: Deep Fake Dataset (Google), Face Forensics++ and manipulated Dataset from
Facebook.
Proposed Methodologies: In this article they have used CNN method and Transferable learning
technique to detect the manipulated videos and images. This transferable learning technique is
used because it transfers the gained knowledge from one source to some other source so the
understanding will be most accurate and efficient. Also, they have auto encoders to reduce the input vector dimension (encoder will do that) and decoder will be used to revive initial
dimensions of given image.
This paper discusses what are the inconsistencies that are performed in recordings or image due to deepfake creation and states a spatiotemporal good model generated by help of LSTM model.
The proposed model mainly finds all the inconsistencies introduced in deepfakes and identifies
original and manipulated videos.
2.6. DEEPFAKE DETECTION BY ANALYSING CONVOLUTIONAL TRACES
Dataset: DFDC Test set, Kaggle .
This paper introduces a new face detection method will be used. The purpose of the Deepfake
video analysis is to develop a novel detection approach that uses an Expectation Maximization
(EM) algorithm to discover a forensics trace hidden in images or videos and feature extraction is completed on a set of local features which then fed to model that’s underlying the
convolutional generative method extracts a group of native options specifically self-addressed
to model. Ad-hoc validations were used to distinguish between fake images made by 5 different realistic architectures of GANs (GDWCT, STARGAN, ATTGAN, STYLEGAN,
STYLEGAN2) and actual images against the CELEBA dataset using experimental tests using
Nave Bayes classifiers.
Results obtained demonstrated the effectiveness and accuracy of the technique in distinguishing
the different architectures and the corresponding generation process.
Dataset: CELEBA
Merits: This method will be beneficial for making Deep fake predictions, which are very
important for forensic investigations because it can not only categorize a image or video as fake
but also forecast the most likely approach utilized for fabrication.. The maximum classification
accuracy was around 90percent when original dataset was compared with all deep fake
generated models obtained.
3. DATA COLLECTION
YouTube Data-set is used for Fake Image Detection 7104 images for 2845 fake images and
4259 real images which contains faces images from different sources on YouTube.This is used
for Fake Image Detection model.Testing is done from https://thispersondoesnotexist.com
• Face Forensics++ could be a forensics data-set consisting of 1000 original video sequences
that are manipulated with Deep-fakes, Face2Face, Face Swap and Neural Textures these four automated face manipulation methods. For this the information has been sourced from 977
different YouTube videos and each video includes mostly frontal face.Sample around 600
videos of both kinds are used in training to have a balanced data-set.
• Celeb-DF this data-set contains high-quality 5,639 different Deep-Fake videos of different celebrities from many industries that one is generated using improved synthesis process.Sample
of 1000 videos are used training the model.
• Deep-fake detection challenge Data-set from Kaggle Competition. From these data-sets
randomly selecting videos to create new data-set for Fake Video Detec tion.Around 4000 sample videos are randomly selected and processed by cropping face landmarks and fed in the
Neural network.Total of 6500 videos are used to train the model with different pre-trained
models Res net ,Inception Res-net ,Efficient-Net and best accuracy prediction is considered.
• Testing data-set consists of popular celebrity deep-fake videos with original ones and high
quality videos and low videos are used to test result of the trained model.
4. SYSTEM ARCHITECTURE
Figure. 1. System Architecture of Deep-Fake Detection.
4.1 PROPOSED APPROACH
• Apply the pre-processing step required for Deep fake detection which basically involves cropping the face images and creating new data-set which are fed to deep learning models.
• We are applying required transformations to our training images that has been divided into two classes that is fake(manipulated) and real.
• Initialize and loading the model (base model for that initial layers are readily available).
• After that we are Loading our base model (MesoNet/ResNet) in the mentioned above step with assigning weights then we obtain new pre-processed data which has 6000 videos from
Face Forensics ++, Celeb DF, Deep-fake detection challenge data-set.
• Adding required extra and custom layers in the model. • For model we are initializing dropout, learning rate such kind of hyper parameters to improve the features. • And
Training of the model with features of eye blinking, facial transformations, head poses, blurring, inconsistency in the intensity of the face, the most important feature is finding
similarity in the light reflections in the eyeballs. • Once we complete training of the model,
we need test the model on them with re-scaling the images in test model.MesoNet model is
used for Fake image detection on YouTube data-set of 7104 images.
ResNet50 and Xception base models were used for implementation of main model for Fake video detection. Inception-ResNet-v2 is a convolutional neural network that is
trained on more than a million images from the Image-Net database.
• Another new hybrid model used is Inception is created to serve the purpose of reducing the computational burden of deep neural nets while ob taining state-of-art performance.While
Inception focuses on computational cost, ResNet focuses on computational accuracy.
• Ensemble method is used to make predication from these base learners where results are stored are meta learners, from these consensus algorithm is implemented to pro vide the best result with best confidence of prediction.
5. MODEL ARCHITECTURE
Figure. 2. Training Flow chart.
Figure. 3. Prediction Flow chart.
EAR takes six points(pi) around the eyes and calculates the absolute area of the horizontal axis
and vertical axis.
Figure. 4. Eye Tracking Count as prepossessing step.
EAR formula used to detect eye blinks, as defined in research. Point’s p1 and p4 refer to the
horizontal axis in the eye area, and the other points refer to the vertical axis. Thus,
EAR is an absolute value of the size calculated through the area of the horizontal and vertical
axes.
EARi = (EARl + EARr)/2 (2)
6. ALGORITHM
• Step 1 : Importing Required Libraries for the model (Tensor-flow and CV)
• Step 2 :Data-set collection of videos with labels. • Step 3 : Pre Processing the data by cropping face region • Step 4 : Split frames of video to images
• Step 5 : Define the model architecture of Meso Net, Res Net50 , Inception-ResNet-v2 base models.
• Step 6 : Convert generated images to arrays and provide as input Neural Net.
• Step 7 : Specifying parameters ,loss function , initializ ers,Adam optimizer
• Step 8 :Splitting image data into training and validation set.
• Step 8 : Training the model for 20 epochs
• Step 9 : Save the trained model
• Step 10:Predict the result of unknown video.
• Step 11:Test the model on high and low videos.
• Step 12 :Perform the metrics, accuracy and loss and time and space complexity
7. METHODOLOGY
Eye blinking can be done in two ways: one is to close one’s eyelids, and the other is process of
lifting one’s closed eyelids. As a result, we used EARi to develop a method for identifying eye blinks .This is how EAR formula helps in determining Eye blink counts and if deep-fake video
goes to this eye tracker it finds that eye blink count will less 10 where we find persistent eye
blink count per count will 15-25 not less than that in normal scenarios.
EAR = ||p2 − p6|| + ||p3 − p5||/2||p1 − p4|| (1)
Also, both frames of left and right eyeball region will be captured and compare the similarity using similarity index and if similarity index in eyeball which captured the light reflections is
low then if it considered as fake video else further it will be moved to further deep learning
algorithms like CNN and LSTM.
Data-set is then loaded to ResNeXt-50 5 Convolution layers then fed into LSTM networks and final detection is made. Here first 5 layers of Convolutional neural network (CNN) for feature
extraction. And also used pretrained models like Meso Net and ResNext50 CNN classifier
which extract the features with accurately detect the frame-level features. Then we are fine-
tuning that network by adding required extra layers and choosing a legitimate learning rate to appropriately combine the gradient descent of the model. For Sequence pre-processing Long
short-term memory (LSTM) will be used. It helps to process the obtained frames in a successive
manner with help of temporal analysis once LSTM model done with its work the data further transferred for model evaluation where we are generating a confusion matrix for model
evaluation purpose and finally, we are loading the trained model, which will predict the input
data is Real or Fake. Prediction of unknown video is done, first preliminary test for fake
detection is done by checking Eye Blink rate count then based on trained model it checks each frames and final prediction is made.
Figure. 5. Prediction of ResNet-LSTM model.
Figure. 6. Model Architecture of Inception-ResNet-v2 base model.
8.MODEL EVALUATION
Models trained are then evaluated by finding accuracy, precision, F1 score and calculating
Confusion Matrix. If the accuracy is high then the model created is suitable for testing, so, exporting the trained model for prediction. Then finally after testing the highest accuracy model
will be used for Deployment of the model. Accuracy of the model is 90.123 percent using
MesoNet for Fake Image detection.Model evaluation is performed using Confusion matrix.
Fig. 7. Model Accuracy and Loss of Pre-trained InceptionResNetV2.
9.TESTING
Pre-trained model of InceptionResNetV2 was tested on popular deep-fake videos
Figure. 8. Table containing popular fake videos and its prediction.
Figure. 9. Confusion Matrix of Fake Image Detection
Figure. 10. Confusion Matrix of Fake Video Detection
10. CONCLUSION
Fake Image Detection using MesoNet produced an accuracy of 90.123 and Fake Video
Detection using ResNet gave Accuracy of 86.39 using InceptionResNetV2 gave Accuracy of
89.54 nearly 90 percent and better model with better testing results. Model testing using
various deep learning techniques is required and there is need to test results on the fake videos and popular fake videos and also fake audio as well which are used to manipulate voices . A way to deal with improve execution is to make a developing refreshed benchmark
informational collection of deep-fakes to approve the continuous advancement of recognition
techniques and using multiple models that give high accuracy and testing it on real-time
applications is one way to improve detection technique performance. Hence, the development of new and more robust methods to deal with the increasingly challenging deep fakes is the
goal of this research paper.
11. ACKNOWLEDGEMENTS
I would like to express my gratitude to Dr. Jayashree R, Department of Computer Science and
Engineering, PES University, for her continuous guidance, assistance, and encouragement
throughout the development of this Capstone Project.I take this opportunity to thank Dr. Shylaja S S, Chairperson, Department of Computer Science and Engineering, PES University,
for all the knowledge and support I have received from the department. I would like to thank
Dr. B.K. Keshavan, Dean of Faculty, PES University for his help.
12. REFERENCES
[1] Badrinarayanan, V., Kendall, A., and Cipolla, R. (2017). SegNet: A deep convolutional encoder-decoder architecture for image segmenta tion. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(12), 2481-2495. [2] ang, W., Hui, C., Chen, Z., Xue, J. H., and Liao, Q. (2019). FV-GAN: Finger vein representation using generative adversarial networks. IEEE Transactions on Information Forensics and Security, 14(9), 2512-2524. [3] Tewari, A., Zollhoefer, M., Bernard, F., Garrido, P., Kim, H., Perez, P., and Theobalt, C. (2020). High-fidelity monocular face reconstruc tion based on an unsupervised model-based face
autoencoder. IEEE Transactions on Pattern Analysis and Machine Intelligence. DOI: 10.1109/TPAMI.2018.2876842. [4] Guo, Y., Jiao, L., Wang, S., Wang, S., and Liu, F. (2018). Fuzzy sparse autoencoder framework for single image per person face recognition. IEEE Transactions on Cybernetics, 48(8), 2402-2415. [5] Liu, F., Jiao, L., and Tang, X. (2019). Task-oriented GAN for PolSAR image classification and clustering. IEEE Transactions on Neural Net works and Learning Systems, 30(9), 2707- 2719. [6] Y. Yorozu, M. Hirano, K. Oka, and Y. Tagawa, “Electron spectroscopy studies on magneto- optical media and plastic substrate interface,” IEEE Transl. J. Magn. Japan, vol. 2, pp. 740–741, August 1987 [Digests 9th Annual Conf. Magnetics Japan, p. 301, 1982]. [7] Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., and Metaxas, D. N. (2019). StackGAN++: Realistic image synthesis with stacked generative adversarial networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(8), 1947-1962. [8] Bloomberg (2018, September 11). How faking videos became easy and why that’s so scary. Available at https://fortune.com/2018/09/11/deep fakes-obama-video/ [9] Chesney, R., and Citron, D. (2019). Deepfakes and the new disinfor mation war: The coming age of post-truth geopolitics. Foreign Affairs, 98, 147. [10] Kaliyar, R. K., Goswami, A., and Narang, P. (2020). Deep fake: improving fake news detection using tensor decomposi tion based deep neural network. Journal of Supercomputing, doi: https://doi.org/10.1007/s11227-020-03294-y. [11] ucker, P. (2019, March 31). The newest AI-enabled weapon: Deep-Faking photos of the earth. Available at https://www.defenseone.com/technology/2019/03/next-phase-ai deep faking- whole-world-and-china-ahead/155944/ [12] Damiani, J. (2019, September 3). A voice deepfake was used to scam a CEO out of 243,000. Available at https://www.forbes.com/sites/jessedamiani/2019/09/03/a-voice deepfake- was-used-to-scam-a-ceo-out-of-243000/ [13] Jafar, M. T., Ababneh, M., Al-Zoube, M., and Elhassan, A. (2020, April). Forensics and analysis of deepfake videos. In The 11th Interna tional Conference on Information and Communication Systems (ICICS) (pp. 053-058). IEEE. [14] Lyu, S. (2020, July). Deepfake detection: current challenges and next steps. In IEEE International Conference on Multimedia and Expo Workshops (ICMEW) (pp. 1-6). IEEE. [15] Shraddha Suratkar,Karan Variyambat,(July 1-3, 2020) Employing Transfer-Learning based CNN architectures to Enhance the Generaliz ability of Deep-fake Detection [16] Akul Mehra(July 1-3, 2020) Deep-fake Detection using Capsule Networks with Long Short-Term Memory Networks
o
- Deep-fake Detection using CNN Ensemble of learners with Transfer Learning Techniques
- Yathish N V1, Dr. Jayashree 2, Jagadish Rathod3, Manu L4 and Prsahant 5
- 12345 Department of Computer Science and Engineering, PES University, Bengaluru, India
- [email protected]
- abstract
- Today, Computer science and its applications are growing rapidly with recent advancement of the technology. Particularly this deep learning algorithms, image processing, fake media creations have gained lot of recognition and creating threat to the pe...
- The proposal of this project is to examine these media files for AI-generated media alterations in Faces and develop new, more robust, and sophisticated approaches to deal with the more difficult deep fakes which gets challenging everyday. The propose...
- keywords
- Deep learning, deep fake, Artificial intelligence, Deep-fake detection, transfer learning, CNN,Inception-ResNet-v2.,ResNet,LSTM
- 1. INTRODUCTION
- The term "Deep-fake" refers to a technology that uses Deep Learning to create fake videos in which one person's face is swapped/morphed onto the face of another. Deep-fakes are the outcome of advanced machine learning and AI techniques used to manipul...
- GANs have the ability to create ultra-realistic fake images and videos. Various powerful machine learning and Artificial intelligence algorithms capable of making highly realistic Deep Fake movies are being used with malicious purpose to deceive peopl...
- 2. PREVIOUS WORK
- 2.1. Deep Learning for Deep Fakes Creation and Detection: A Survey
- 2.2. Deep fake Detection in Media Files - Audios, Images and Videos
- 2.3Video Detection Using Recurrent Neural Networks (RNN)
- 11. Acknowledgements
- 12. References
- Authors
- Yathish N V
- Dr. Jayashree R.
- Prashant
- Jagadish Rathod
- Manu L
- o
- o (1)