3 papers
cs.SD2024
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
Zhiyong Wang, Ruibo Fu, Zhengqi Wen +10
Speech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and in…
cs.SD2024
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
Xin Qi, Ruibo Fu, Zhengqi Wen +12
In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have…
eess.AS2020
Deep Attention Fusion Feature for Speech Separation with End-to-End Post-filter Method
Cunhang Fan, Jianhua Tao, Bin Liu +3
In this paper, we propose an end-to-end post-filter method with deep attention fusion features for monaural speaker-independent speech separation. At first, a time-frequency domain…