7 papers
DFALLM: Achieving Generalizable Multitask Deepfake Detection by Optimizing Audio LLM Components
Yupei Li, Li Wang, Yuxiang Wang +5
Audio deepfake detection has recently garnered public concern due to its implications for security and reliability. Traditional deep learning methods have been widely applied to th…
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
Xueyao Zhang, Chaoren Wang, Huan Liao +8
Aligning large generative models with human feedback is a critical challenge. In speech synthesis, this is particularly pronounced due to the lack of a large-scale human preference…
Over-the-Air Adversarial Attack Detection: from Datasets to Defenses
Li Wang, Xiaoyan Lei, Haorui He +3
Automatic Speaker Verification (ASV) systems can be used for voice-enabled applications for identity verification. However, recent studies have exposed these systems' vulnerabiliti…
Audio Deepfake Verification
Li Wang, Junyi Ao, Linyong Gan +3
With the rapid development of deepfake technology, simply making a binary judgment of true or false on audio is no longer sufficient to meet practical needs. Accurately determining…
Overview of the Amphion Toolkit (v0.2)
Jiaqi Li, Xueyao Zhang, Yuancheng Wang +9
Amphion is an open-source toolkit for Audio, Music, and Speech Generation, designed to lower the entry barrier for junior researchers and engineers in these fields. It provides a v…
Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning
Haorui He, Yuchen Song, Yuancheng Wang +6
The effectiveness of one-shot voice conversion (VC) decreases in real-world scenarios where reference speeches, which are often sourced from the internet, contain various disturban…