10 - read paper and provide feedback/review

profilePXP
Paper11-Copy.pdf

DEEP-FAKE DETECTION USING CNN ENSEMBLE OF

LEARNERS WITH TRANSFER LEARNING

TECHNIQUES

ABSTRACT

Today, Computer science and its applications are growing rapidly with recent advancement of

the technology. Particularly this deep learning algorithms, image processing, fake media creations

have gained lot of recognition and creating threat to the people nowadays. Deep Fake Creation is one

amongst the large threats to the authenticity and confidential of online information. These Deep-fake will

be used for malicious intents like phishing scams and fake news. So, building this Detection system

may well be solution for fraud detection for prevention of widespread attention of fake news,

videos, footage and pictures.

The proposal of this project is to examine these media files for AI-generated media alterations in

Faces and develop new, more robust, and sophisticated approaches to deal with the more difficult

deep fakes which gets challenging everyday. The proposed architecture uses Ensemble of CNNs

containing different base learners that analyses individual face feature like mouth,eye,facial landmarks

that uses pre-trained models like ResNet, Inception-ResNet-v2 ,XceptionNet ,Efficient-Net and final

classifier predicts the best result from the learners .Experiments were conducted on data-sets that

contain both fake and real videos and also with custom built data-set of World Politicians and

Celebrities.Performance of Ensemble method has promising results when compared to other methods.

KEYWORDS

Deep learning, deep fake, Artificial intelligence, Deep-fake detection, transfer learning,

CNN,Inception-ResNet-v2.,ResNet,LSTM

1. INTRODUCTION The term "Deep-fake" refers to a technology that uses Deep Learning to create fake videos

in which one person's face is swapped/morphed onto the face of another. Deep-fakes are the

outcome of advanced machine learning and AI techniques used to manipulate or generate visual

and audio content with a high potential to deceive while imitating content. Deep learning

is primarily responsible for the creation of deep-fakes, which entails training generative

neural network designs such as auto encoders or generative adversarial networks (GANs).

GANs have the ability to create ultra-realistic fake images and videos. Various powerful

machine learning and Artificial intelligence algorithms capable of making highly realistic

Deep Fake movies are being used with malicious purpose to deceive people, harass, blackmail

women, and cause political instability by propagating false news and hostile propaganda, which

could lead to social, political, and economic instability and outbursts with terrible

consequences. Therefore, Deep fake videos pose a severe threat to personal safety, affecting not

just public images but also regular people, as well as national security, demanding the

development of automated methods to detect deep fake movies.. There are numerous public

implementations and usage of Deep fake out there, most notably Fake-App and the face swap

GitHub. Some of the Deep Fake tools are Face Forensics++, Face swap-GAN, DeepFaceLab,

Dfaker, etc. Sentinel is a start-up company

working on fake media content.\\Deep Fake Detection system has lot of applications in Forensics

for investigation of frauds, scams and help in prevention of threats, bullying, blackmails and

prevention of crime. It also helps in prevention of widespread attention of Fake news, videos,

footage and images. It reduces social abuse, harassment, promotes privacy and security of global leaders. So, finding the reality in digital domain so has become progressively essential. In fact,

when it comes to detecting deep-fake is even more challenging.\\Lately in 2019, Facebook Inc.

teaming up with Microsoft has been released the Deep-fake Detection Challenge to catalyze for

more improvement and development with better accuracy in detecting and preventing

manipulated data or deep-fakes in Kaggle and used YouTube data-set, yet remarked the challenge

is unsolved. But recent advancements in 2020, Deep Fake Detection has gained lot of popularity and lot of research have been going and quite become trending these days.Sentinel is a start-up

company working on fake media content.\\This Graph taken from https://app.dimensions.ai/

shows that at the end of August 2020 it displays that the number of papers for deep-fakes has been

increased exponentially in recent years

2. PREVIOUS WORK

2.1. DEEP LEARNING FOR DEEP FAKES CREATION AND DETECTION: A SURVEY

This research paper provides a complete survey for Deep Fake Creation and Detection which

tells about recent works, algorithms used, data-set used performance and application by various researchers. Deep Fake Detection methods are still in the initial stage and there is more

upcoming technologies in the future to detect deep-fakes. Then to summarize the table of

detection methods is displayed which uses different methods like CNN LSTM, VGG19, ResNet50, MesoNet, KNN, SVM, CFFN and many more.

2.2. DEEP FAKE DETECTION IN MEDIA FILES - AUDIOS, IMAGES AND VIDEOS

This study proposes a hybrid strategy that combines image processing and the CNN network.

There are four sections to this methodology. The first is the Data Preparation step, which

converts input samples to visual samples. The second step is the Data Enhancement section, which removes the image’s noise component. Following that is the CNN model, which is used

for training and testing to build the detection model, and finally the detection portion, which

does real-time detection against various media files from multiple platforms. Data-set: Deep-

fake Timit, Forensics++, VidTimit data-set

This paper presents a deep-fake image creation by using GAN network and there we have two networks one is generator and another one is discriminator. Here generator part of the GAN

creates fake videos and then the discriminator detects it.

2.3VIDEO DETECTION USING RECURRENT NEURAL NETWORKS (RNN)

In this research paper, they have introduced a new technique to automatically detect manipulated

videos. For extracting frame-level features they have used a CNN method. These features which

extracted from CNN method are passed to RNN that helps for classifying the given image/video has been manipulated or not. For dimensionality reduction autoencoder has been applied, LSTM

(Long short-term memory) are the one of the particular sub-types of RNN method, LSTMs are

used for temporal sequence analysis. And finally produce and probability likelihood estimation of the sequence being either a manipulated video or a real video.

Proposed Methodology: Methodologies which author used in this paper is Recurrent Neural

Network (RNN) for detect ing deepfake and for extracting frame-level features from every frame

in the video they have used CNN model. The output/features of CNN model used to train a RNN

(Recurrent Neural Network) that will learn to predict if a video/image manipulated or not. Autoencoders have been applied for dimensionality reduction and compactly represent ing the

videos/images. for temporal sequence analysis LSTM method was used.

2.4. DEEPFAKE DETECTION USING TRANSFER-LEARNING BASED CNN ARCHITECTURE

Here the authors introduce a way to detect a manipulated videos by employing a Convolutional

neural network (CNN) and using the Transferable Learning method. The technique used during this paper that one implements a method or architecture includes CNN for extracting features

from each frame of the video for training a model (binary classifier) that will learn efficiently to

differentiate the important and pretend videos. the strategy is evaluated against an extensively more set of manipulated videos those are collected from many different datasets. for generating

Deep Fakes Autoencoder method was used. Nowadays creating or modifying models for every

specific application that may be extremely difficult in terms of cost and also manpower. during this reasonably situation Transfer Learning would be helpful it’ll offer a promising solution to

such problem. the main intension of using this transferable learning method is to transfer the

knowledge which is gained from one source to other source so the understanding will be more

accurate and efficient. By using this application, we are able to build a comparatively good model for deep fake detection.

2.5. DEEPFAKE DETECTION USING CAPSULE NETWORKS WITH LONG SHORT TERM MEMORY

NETWORKS

Dataset: Deep Fake Dataset (Google), Face Forensics++ and manipulated Dataset from

Facebook.

Proposed Methodologies: In this article they have used CNN method and Transferable learning

technique to detect the manipulated videos and images. This transferable learning technique is

used because it transfers the gained knowledge from one source to some other source so the

understanding will be most accurate and efficient. Also, they have auto encoders to reduce the input vector dimension (encoder will do that) and decoder will be used to revive initial

dimensions of given image.

This paper discusses what are the inconsistencies that are performed in recordings or image due to deepfake creation and states a spatiotemporal good model generated by help of LSTM model.

The proposed model mainly finds all the inconsistencies introduced in deepfakes and identifies

original and manipulated videos.

2.6. DEEPFAKE DETECTION BY ANALYSING CONVOLUTIONAL TRACES

Dataset: DFDC Test set, Kaggle .

This paper introduces a new face detection method will be used. The purpose of the Deepfake

video analysis is to develop a novel detection approach that uses an Expectation Maximization

(EM) algorithm to discover a forensics trace hidden in images or videos and feature extraction is completed on a set of local features which then fed to model that’s underlying the

convolutional generative method extracts a group of native options specifically self-addressed

to model. Ad-hoc validations were used to distinguish between fake images made by 5 different realistic architectures of GANs (GDWCT, STARGAN, ATTGAN, STYLEGAN,

STYLEGAN2) and actual images against the CELEBA dataset using experimental tests using

Nave Bayes classifiers.

Results obtained demonstrated the effectiveness and accuracy of the technique in distinguishing

the different architectures and the corresponding generation process.

Dataset: CELEBA

Merits: This method will be beneficial for making Deep fake predictions, which are very

important for forensic investigations because it can not only categorize a image or video as fake

but also forecast the most likely approach utilized for fabrication.. The maximum classification

accuracy was around 90percent when original dataset was compared with all deep fake

generated models obtained.

3. DATA COLLECTION

YouTube Data-set is used for Fake Image Detection 7104 images for 2845 fake images and

4259 real images which contains faces images from different sources on YouTube.This is used

for Fake Image Detection model.Testing is done from https://thispersondoesnotexist.com

• Face Forensics++ could be a forensics data-set consisting of 1000 original video sequences

that are manipulated with Deep-fakes, Face2Face, Face Swap and Neural Textures these four automated face manipulation methods. For this the information has been sourced from 977

different YouTube videos and each video includes mostly frontal face.Sample around 600

videos of both kinds are used in training to have a balanced data-set.

• Celeb-DF this data-set contains high-quality 5,639 different Deep-Fake videos of different celebrities from many industries that one is generated using improved synthesis process.Sample

of 1000 videos are used training the model.

• Deep-fake detection challenge Data-set from Kaggle Competition. From these data-sets

randomly selecting videos to create new data-set for Fake Video Detec tion.Around 4000 sample videos are randomly selected and processed by cropping face landmarks and fed in the

Neural network.Total of 6500 videos are used to train the model with different pre-trained

models Res net ,Inception Res-net ,Efficient-Net and best accuracy prediction is considered.

• Testing data-set consists of popular celebrity deep-fake videos with original ones and high

quality videos and low videos are used to test result of the trained model.

4. SYSTEM ARCHITECTURE

Figure. 1. System Architecture of Deep-Fake Detection.

4.1 PROPOSED APPROACH

• Apply the pre-processing step required for Deep fake detection which basically involves cropping the face images and creating new data-set which are fed to deep learning models.

• We are applying required transformations to our training images that has been divided into two classes that is fake(manipulated) and real.

• Initialize and loading the model (base model for that initial layers are readily available).

• After that we are Loading our base model (MesoNet/ResNet) in the mentioned above step with assigning weights then we obtain new pre-processed data which has 6000 videos from

Face Forensics ++, Celeb DF, Deep-fake detection challenge data-set.

• Adding required extra and custom layers in the model. • For model we are initializing dropout, learning rate such kind of hyper parameters to improve the features. • And

Training of the model with features of eye blinking, facial transformations, head poses, blurring, inconsistency in the intensity of the face, the most important feature is finding

similarity in the light reflections in the eyeballs. • Once we complete training of the model,

we need test the model on them with re-scaling the images in test model.MesoNet model is

used for Fake image detection on YouTube data-set of 7104 images.

ResNet50 and Xception base models were used for implementation of main model for Fake video detection. Inception-ResNet-v2 is a convolutional neural network that is

trained on more than a million images from the Image-Net database.

• Another new hybrid model used is Inception is created to serve the purpose of reducing the computational burden of deep neural nets while ob taining state-of-art performance.While

Inception focuses on computational cost, ResNet focuses on computational accuracy.

• Ensemble method is used to make predication from these base learners where results are stored are meta learners, from these consensus algorithm is implemented to pro vide the best result with best confidence of prediction.

5. MODEL ARCHITECTURE

Figure. 2. Training Flow chart.

Figure. 3. Prediction Flow chart.

EAR takes six points(pi) around the eyes and calculates the absolute area of the horizontal axis

and vertical axis.

Figure. 4. Eye Tracking Count as prepossessing step.

EAR formula used to detect eye blinks, as defined in research. Point’s p1 and p4 refer to the

horizontal axis in the eye area, and the other points refer to the vertical axis. Thus,

EAR is an absolute value of the size calculated through the area of the horizontal and vertical

axes.

EARi = (EARl + EARr)/2 (2)

6. ALGORITHM

• Step 1 : Importing Required Libraries for the model (Tensor-flow and CV)

• Step 2 :Data-set collection of videos with labels. • Step 3 : Pre Processing the data by cropping face region • Step 4 : Split frames of video to images

• Step 5 : Define the model architecture of Meso Net, Res Net50 , Inception-ResNet-v2 base models.

• Step 6 : Convert generated images to arrays and provide as input Neural Net.

• Step 7 : Specifying parameters ,loss function , initializ ers,Adam optimizer

• Step 8 :Splitting image data into training and validation set.

• Step 8 : Training the model for 20 epochs

• Step 9 : Save the trained model

• Step 10:Predict the result of unknown video.

• Step 11:Test the model on high and low videos.

• Step 12 :Perform the metrics, accuracy and loss and time and space complexity

7. METHODOLOGY

Eye blinking can be done in two ways: one is to close one’s eyelids, and the other is process of

lifting one’s closed eyelids. As a result, we used EARi to develop a method for identifying eye blinks .This is how EAR formula helps in determining Eye blink counts and if deep-fake video

goes to this eye tracker it finds that eye blink count will less 10 where we find persistent eye

blink count per count will 15-25 not less than that in normal scenarios.

EAR = ||p2 − p6|| + ||p3 − p5||/2||p1 − p4|| (1)

Also, both frames of left and right eyeball region will be captured and compare the similarity using similarity index and if similarity index in eyeball which captured the light reflections is

low then if it considered as fake video else further it will be moved to further deep learning

algorithms like CNN and LSTM.

Data-set is then loaded to ResNeXt-50 5 Convolution layers then fed into LSTM networks and final detection is made. Here first 5 layers of Convolutional neural network (CNN) for feature

extraction. And also used pretrained models like Meso Net and ResNext50 CNN classifier

which extract the features with accurately detect the frame-level features. Then we are fine-

tuning that network by adding required extra layers and choosing a legitimate learning rate to appropriately combine the gradient descent of the model. For Sequence pre-processing Long

short-term memory (LSTM) will be used. It helps to process the obtained frames in a successive

manner with help of temporal analysis once LSTM model done with its work the data further transferred for model evaluation where we are generating a confusion matrix for model

evaluation purpose and finally, we are loading the trained model, which will predict the input

data is Real or Fake. Prediction of unknown video is done, first preliminary test for fake

detection is done by checking Eye Blink rate count then based on trained model it checks each frames and final prediction is made.

Figure. 5. Prediction of ResNet-LSTM model.

Figure. 6. Model Architecture of Inception-ResNet-v2 base model.

8.MODEL EVALUATION

Models trained are then evaluated by finding accuracy, precision, F1 score and calculating

Confusion Matrix. If the accuracy is high then the model created is suitable for testing, so, exporting the trained model for prediction. Then finally after testing the highest accuracy model

will be used for Deployment of the model. Accuracy of the model is 90.123 percent using

MesoNet for Fake Image detection.Model evaluation is performed using Confusion matrix.

Fig. 7. Model Accuracy and Loss of Pre-trained InceptionResNetV2.

9.TESTING

Pre-trained model of InceptionResNetV2 was tested on popular deep-fake videos

Figure. 8. Table containing popular fake videos and its prediction.

Figure. 9. Confusion Matrix of Fake Image Detection

Figure. 10. Confusion Matrix of Fake Video Detection

10. CONCLUSION

Fake Image Detection using MesoNet produced an accuracy of 90.123 and Fake Video

Detection using ResNet gave Accuracy of 86.39 using InceptionResNetV2 gave Accuracy of

89.54 nearly 90 percent and better model with better testing results. Model testing using

various deep learning techniques is required and there is need to test results on the fake videos and popular fake videos and also fake audio as well which are used to manipulate voices . A way to deal with improve execution is to make a developing refreshed benchmark

informational collection of deep-fakes to approve the continuous advancement of recognition

techniques and using multiple models that give high accuracy and testing it on real-time

applications is one way to improve detection technique performance. Hence, the development of new and more robust methods to deal with the increasingly challenging deep fakes is the

goal of this research paper.

11. ACKNOWLEDGEMENTS

I would like to express my gratitude to Dr. Jayashree R, Department of Computer Science and

Engineering, PES University, for her continuous guidance, assistance, and encouragement

throughout the development of this Capstone Project.I take this opportunity to thank Dr. Shylaja S S, Chairperson, Department of Computer Science and Engineering, PES University,

for all the knowledge and support I have received from the department. I would like to thank

Dr. B.K. Keshavan, Dean of Faculty, PES University for his help.

12. REFERENCES

[1] Badrinarayanan, V., Kendall, A., and Cipolla, R. (2017). SegNet: A deep convolutional encoder-decoder architecture for image segmenta tion. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(12), 2481-2495. [2] ang, W., Hui, C., Chen, Z., Xue, J. H., and Liao, Q. (2019). FV-GAN: Finger vein representation using generative adversarial networks. IEEE Transactions on Information Forensics and Security, 14(9), 2512-2524. [3] Tewari, A., Zollhoefer, M., Bernard, F., Garrido, P., Kim, H., Perez, P., and Theobalt, C. (2020). High-fidelity monocular face reconstruc tion based on an unsupervised model-based face

autoencoder. IEEE Transactions on Pattern Analysis and Machine Intelligence. DOI: 10.1109/TPAMI.2018.2876842. [4] Guo, Y., Jiao, L., Wang, S., Wang, S., and Liu, F. (2018). Fuzzy sparse autoencoder framework for single image per person face recognition. IEEE Transactions on Cybernetics, 48(8), 2402-2415. [5] Liu, F., Jiao, L., and Tang, X. (2019). Task-oriented GAN for PolSAR image classification and clustering. IEEE Transactions on Neural Net works and Learning Systems, 30(9), 2707- 2719. [6] Y. Yorozu, M. Hirano, K. Oka, and Y. Tagawa, “Electron spectroscopy studies on magneto- optical media and plastic substrate interface,” IEEE Transl. J. Magn. Japan, vol. 2, pp. 740–741, August 1987 [Digests 9th Annual Conf. Magnetics Japan, p. 301, 1982]. [7] Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., and Metaxas, D. N. (2019). StackGAN++: Realistic image synthesis with stacked generative adversarial networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41(8), 1947-1962. [8] Bloomberg (2018, September 11). How faking videos became easy and why that’s so scary. Available at https://fortune.com/2018/09/11/deep fakes-obama-video/ [9] Chesney, R., and Citron, D. (2019). Deepfakes and the new disinfor mation war: The coming age of post-truth geopolitics. Foreign Affairs, 98, 147. [10] Kaliyar, R. K., Goswami, A., and Narang, P. (2020). Deep fake: improving fake news detection using tensor decomposi tion based deep neural network. Journal of Supercomputing, doi: https://doi.org/10.1007/s11227-020-03294-y. [11] ucker, P. (2019, March 31). The newest AI-enabled weapon: Deep-Faking photos of the earth. Available at https://www.defenseone.com/technology/2019/03/next-phase-ai deep faking- whole-world-and-china-ahead/155944/ [12] Damiani, J. (2019, September 3). A voice deepfake was used to scam a CEO out of 243,000. Available at https://www.forbes.com/sites/jessedamiani/2019/09/03/a-voice deepfake- was-used-to-scam-a-ceo-out-of-243000/ [13] Jafar, M. T., Ababneh, M., Al-Zoube, M., and Elhassan, A. (2020, April). Forensics and analysis of deepfake videos. In The 11th Interna tional Conference on Information and Communication Systems (ICICS) (pp. 053-058). IEEE. [14] Lyu, S. (2020, July). Deepfake detection: current challenges and next steps. In IEEE International Conference on Multimedia and Expo Workshops (ICMEW) (pp. 1-6). IEEE. [15] Shraddha Suratkar,Karan Variyambat,(July 1-3, 2020) Employing Transfer-Learning based CNN architectures to Enhance the Generaliz ability of Deep-fake Detection [16] Akul Mehra(July 1-3, 2020) Deep-fake Detection using Capsule Networks with Long Short-Term Memory Networks

o

  • Deep-fake Detection using CNN Ensemble of learners with Transfer Learning Techniques
  • Yathish N V1, Dr. Jayashree 2, Jagadish Rathod3, Manu L4 and Prsahant 5
  • 12345 Department of Computer Science and Engineering, PES University, Bengaluru, India
  • [email protected]
  • abstract
  • Today, Computer science and its applications are growing rapidly with recent advancement of the technology. Particularly this deep learning algorithms, image processing, fake media creations have gained lot of recognition and creating threat to the pe...
  • The proposal of this project is to examine these media files for AI-generated media alterations in Faces and develop new, more robust, and sophisticated approaches to deal with the more difficult deep fakes which gets challenging everyday. The propose...
  • keywords
    • Deep learning, deep fake, Artificial intelligence, Deep-fake detection, transfer learning, CNN,Inception-ResNet-v2.,ResNet,LSTM
  • 1. INTRODUCTION
  • The term "Deep-fake" refers to a technology that uses Deep Learning to create fake videos in which one person's face is swapped/morphed onto the face of another. Deep-fakes are the outcome of advanced machine learning and AI techniques used to manipul...
  • GANs have the ability to create ultra-realistic fake images and videos. Various powerful machine learning and Artificial intelligence algorithms capable of making highly realistic Deep Fake movies are being used with malicious purpose to deceive peopl...
  • 2. PREVIOUS WORK
  • 2.1. Deep Learning for Deep Fakes Creation and Detection: A Survey
  • 2.2. Deep fake Detection in Media Files - Audios, Images and Videos
  • 2.3Video Detection Using Recurrent Neural Networks (RNN)
  • 11. Acknowledgements
  • 12. References
  • Authors
  • Yathish N V
  • Dr. Jayashree R.
  • Prashant
  • Jagadish Rathod
  • Manu L
  • o
  • o (1)