8 papers · 1 filter
Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling
Ye-Xin Lu, Xin Wang, Yang Ai +3
Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or enta…
Does Fine-tuning by Reinforcement Learning Improve Generalization in Binary Speech Deepfake Detection?
Xin Wang, Ge Wanying, Junichi Yamagishi
Building speech deepfake detection models that are generalizable to unseen attacks remains a challenging problem. Although the field has shifted toward a pre-training and fine-tuni…
Human perception of audio deepfakes: the role of language and speaking style
Eugenia San Segundo, Aurora López-Jareño, Xin Wang +1
Audio deepfakes have reached a level of realism that makes it increasingly difficult to distinguish between human and artificial voices, which poses risks such as identity theft or…
Post-training for Deepfake Speech Detection
Wanying Ge, Xin Wang, Xuechen Liu +1
We introduce a post-training approach that adapts self-supervised learning (SSL) models for deepfake speech detection by bridging the gap between general pre-training and domain-sp…
FakeMark: Deepfake Speech Attribution With Watermarked Artifacts
Wanying Ge, Xin Wang, Junichi Yamagishi
Deepfake speech attribution remains challenging for existing solutions. Classifier-based solutions often fail to generalize to domain-shifted samples, and watermarking-based soluti…
Towards Data Drift Monitoring for Speech Deepfake Detection in the context of MLOps
Xin Wang, Wanying Ge, Junichi Yamagishi
When being delivered in applications or services on the cloud, static speech deepfake detectors that are not updated will become vulnerable to newly created speech deepfake attacks…