3 papers
cs.SD2026
Diffusion Reconstruction towards Generalizable Audio Deepfake Detection
Bo Cheng, Songjun Cao, Xiaoming Zhang +3
Achieving robust generalization against unseen attacks remains a challenge in Audio Deepfake Detection (ADD), driven by the rapid evolution of generative models. To address this, w…
cs.SD2025
FreeCodec: A disentangled neural speech codec with fewer tokens
Youqiang Zheng, Weiping Tu, Yueteng Kang +5
Neural speech codecs have gained great attention for their outstanding reconstruction with discrete token representations. It is a crucial component in generative tasks such as spe…
cs.SD2025
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis
Rui Niu, Weihao Wu, Jie Chen +2
Controllable speech synthesis aims to control the style of generated speech using reference input, which can be of various modalities. Existing face-based methods struggle with rob…