16 papers
Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling
Ye-Xin Lu, Xin Wang, Yang Ai +3
Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or enta…
Toward Interpretable Speech Deepfake Detection using Artifact-Specific Experts and Calibrated Detection Scores
Viola Negroni, Xin Wang, Wanying Ge +3
In this work, we propose an interpretable framework for speech deepfake detection based on artifact-specific expert models. Rather than relying on black-box decisions, the framewor…
Dynamic Linear Attention
Xin Wang, Hui Shen, Boyuan Zheng +7
The scalability of Large Language Models (LLMs) to long contexts is fundamentally constrained by the quadratic complexity of standard attention, motivating the adoption of linear a…
Self Voice Conversion as an Attack against Neural Audio Watermarking
Yigitcan Ãzer, Wanying Ge, Zhe Zhang +2
Audio watermarking embeds auxiliary information into speech while maintaining speaker identity, linguistic content, and perceptual quality. Although recent advances in neural and d…
Does Fine-tuning by Reinforcement Learning Improve Generalization in Binary Speech Deepfake Detection?
Xin Wang, Ge Wanying, Junichi Yamagishi
Building speech deepfake detection models that are generalizable to unseen attacks remains a challenging problem. Although the field has shifted toward a pre-training and fine-tuni…
Zero-Day Audio DeepFake Detection via Retrieval Augmentation and Profile Matching
Xuechen Liu, Xin Wang, Junichi Yamagishi
Modern audio deepfake detectors built on foundation models and large training datasets achieve promising detection performance. However, they struggle with zero-day attacks, where…