Deepfake: Definitions, Performance Metrics and Standards, Datasets and Benchmarks, and a Meta-Review
arXiv:2208.10913 · doi:10.3389/fdata.2024.1400024
Abstract
Recent advancements in AI, especially deep learning, have contributed to a significant increase in the creation of new realistic-looking synthetic media (video, image, and audio) and manipulation of existing media, which has led to the creation of the new term ``deepfake''. Based on both the research literature and resources in English and in Chinese, this paper gives a comprehensive overview of deepfake, covering multiple important aspects of this emerging concept, including 1) different definitions, 2) commonly used performance metrics and standards, and 3) deepfake-related datasets, challenges, competitions and benchmarks. In addition, the paper also reports a meta-review of 12 selected deepfake-related survey papers published in 2020 and 2021, focusing not only on the mentioned aspects, but also on the analysis of key challenges and recommendations. We believe that this paper is the most comprehensive review of deepfake in terms of aspects covered, and the first one covering both the English and Chinese literature and sources.
31 pages; study completed by end of July 2021
References in corpus (15)
- WaveNet: A Generative Model for Raw Audio
- Robust Real-World Image Super-Resolution against Adversarial Attacks
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection
- Unmasking DeepFakes with simple Features
- WaveFake: A Data Set to Facilitate Audio Deepfake Detection
- WaveCycleGAN2: Time-domain Neural Post-filter for Speech Waveform Generation
- CHEAT: A Large-scale Dataset for Detecting ChatGPT-writtEn AbsTracts
- Robustness and Generalizability of Deepfake Detection: A Study with Diffusion Models
- X-IQE: eXplainable Image Quality Evaluation for Text-to-Image Generation with Visual Large Language Models
- HC3 Plus: A Semantic-Invariant Human ChatGPT Comparison Corpus
- MLAAD: The Multi-Language Audio Anti-Spoofing Dataset
- EmoFake: An Initial Dataset for Emotion Fake Audio Detection
- AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset
- Towards A Better Metric for Text-to-Video Generation