7 papers
Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
Jun Xue, Zhuolin Yi, Yanzhen Ren +6
Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to controlled additive noise scenarios, lac…
Exploring the Scale and Diversity of Speech Anti-spoofing Datasets: Experiments and Analysis
Zhuolin Yi, Jun Xue, Yanzhen Ren +5
The scale of speech anti-spoofing datasets has grown exponentially over the past decade, driven by the assumption that larger data leads to better performance. However, it remains…
Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
Jun Xue, Tong Zhang, Zhuolin Yi +4
The rapid advancement of generative AI has made audio deepfakes increasingly indistinguishable from authentic human vocals, posing significant threats to persons-of-interest (POI)…
RTCFake: Speech Deepfake Detection in Real-Time Communication
Jun Xue, Zhuolin Yi, Yihuan Huang +6
With the rapid advancement of speech generation technologies, the threat posed by speech deepfakes in real-time communication (RTC) scenarios has intensified. However, existing det…
When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse
Yihuan Huang, Jun Xue, Liu Jiajun +5
Audio-Visual Speech Recognition (AVSR) has achieved remarkable progress in offline conditions, yet its robustness in real-world video conferencing (VC) remains largely unexplored.…
How Well Do Current Speech Deepfake Detection Methods Generalize to the Real World?
Daixian Li, Jun Xue, Yanzhen Ren +4
Recent advances in speech synthesis and voice conversion have greatly improved the naturalness and authenticity of generated audio. Meanwhile, evolving encoding, compression, and t…