A Survey on Speech Deepfake Detection
arXiv:2404.13914 · doi:10.1145/3714458
Abstract
The availability of smart devices leads to an exponential increase in multimedia content. However, advancements in deep learning have also enabled the creation of highly sophisticated Deepfake content, including speech Deepfakes, which pose a serious threat by generating realistic voices and spreading misinformation. To combat this, numerous challenges have been organized to advance speech Deepfake detection techniques. In this survey, we systematically analyze more than 200 papers published up to March 2024. We provide a comprehensive review of each component in the detection pipeline, including model architectures, optimization techniques, generalizability, evaluation metrics, performance comparisons, available datasets, and open source availability. For each aspect, we assess recent progress and discuss ongoing challenges. In addition, we explore emerging topics such as partial Deepfake detection, cross-dataset evaluation, and defences against adversarial attacks, while suggesting promising research directions. This survey not only identifies the current state of the art to establish strong baselines for future experiments but also offers clear guidance for researchers aiming to enhance speech Deepfake detection systems.
38 pages. This paper has been accepted by ACM Computing Surveys
References in corpus (29)
- Towards End-to-End Synthetic Speech Detection
- Scaling Speech Technology to 1,000+ Languages
- Tandem Assessment of Spoofing Countermeasures and Automatic Speaker Verification: Fundamentals
- Audio Deepfake Detection Based on a Combination of F0 Information and Real Plus Imaginary Spectrogram Features
- MelGAN-VC: Voice Conversion and Audio Style Transfer on arbitrarily long samples using Spectrograms
- An Initial Investigation for Detecting Vocoder Fingerprints of Fake Audio
- Synthesized Speech Detection Using Convolutional Transformer-Based Spectrogram Analysis
- Real-time Detection of AI-Generated Speech for DeepFake Voice Conversion
- End-to-End Spectro-Temporal Graph Attention Networks for Speaker Verification Anti-Spoofing and Speech Deepfake Detection
- MLAAD: The Multi-Language Audio Anti-Spoofing Dataset
- ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale
- Low-rank Adaptation Method for Wav2vec2-based Fake Audio Detection
- ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild
- An explainability study of the constant Q cepstral coefficient spoofing countermeasure for automatic speaker verification
- TranssionADD: A multi-frame reinforcement based sequence tagging model for audio deepfake detection
- TIMIT-TTS: a Text-to-Speech Dataset for Multimodal Synthetic Media Detection
- SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
- The DKU-DUKEECE System for the Manipulation Region Location Task of ADD 2023
- CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
- Can large-scale vocoded spoofed data improve speech spoofing countermeasure with a self-supervised front end?
- Can spoofing countermeasure and speaker verification systems be jointly optimised?
- Characterizing the temporal dynamics of universal speech representations for generalizable deepfake detection
- Attentive activation function for improving end-to-end spoofing countermeasure systems
- Every Breath You Don't Take: Deepfake Speech Detection Using Breath
- The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio
- Adaptive Fake Audio Detection with Low-Rank Model Squeezing
- Multi-perspective Information Fusion Res2Net with RandomSpecmix for Fake Speech Detection
- Multi-Task Learning in Utterance-Level and Segmental-Level Spoof Detection
- Baseline Systems for the First Spoofing-Aware Speaker Verification Challenge: Score and Embedding Fusion
Cited by in corpus (4)
- Interpreting Multi-Branch Anti-Spoofing Architectures: Correlating Internal Strategy with Empirical Performance
- Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction
- Audio Deepfake Detection at the First Greeting: "Hi!"
- Frame-level Temporal Difference Learning for Partial Deepfake Speech Detection