Combining EfficientNet and Vision Transformers for Video Deepfake Detection
arXiv:2107.02612 · doi:10.1007/978-3-031-06433-3_19
Abstract
Deepfakes are the result of digital manipulation to forge realistic yet fake imagery. With the astonishing advances in deep generative models, fake images or videos are nowadays obtained using variational autoencoders (VAEs) or Generative Adversarial Networks (GANs). These technologies are becoming more accessible and accurate, resulting in fake videos that are very difficult to be detected. Traditionally, Convolutional Neural Networks (CNNs) have been used to perform video deepfake detection, with the best results obtained using methods based on EfficientNet B7. In this study, we focus on video deep fake detection on faces, given that most methods are becoming extremely accurate in the generation of realistic human faces. Specifically, we combine various types of Vision Transformers with a convolutional EfficientNet B0 used as a feature extractor, obtaining comparable results with some very recent methods that use Vision Transformers. Differently from the state-of-the-art approaches, we use neither distillation nor ensemble methods. Furthermore, we present a straightforward inference procedure based on a simple voting scheme for handling multiple faces in the same video shot. The best model achieved an AUC of 0.951 and an F1 score of 88.0%, very close to the state-of-the-art on the DeepFake Detection Challenge (DFDC).
References in corpus (7)
- Generative Adversarial Networks
- DeepFakes: a New Threat to Face Recognition? Assessment and Detection
- Deepfake Video Detection Using Convolutional Vision Transformer
- Deepfake Detection using Spatiotemporal Convolutional Networks
- Efficient Document Re-Ranking for Transformers by Precomputing Term Representations
- Deepfake Detection Scheme Based on Vision Transformer and Distillation
- Solving the Same-Different Task with Convolutional Neural Networks
Cited by in corpus (6)
- Deep Convolutional Pooling Transformer for Deepfake Detection
- A Review of Deep Learning-based Approaches for Deepfake Content Detection
- Cross-Forgery Analysis of Vision Transformers and CNNs for Deepfake Image Detection
- UMMAFormer: A Universal Multimodal-adaptive Transformer Framework for Temporal Forgery Localization
- Enhancing Deepfake Detection using SE Block Attention with CNN
- Towards Exploring Fairness in Visual Transformer based Natural and GAN Image Detection Systems