1 citations · 1 across the 3 of their papers we have counts for
Showing eess.ASShow all
3 papers · 1 filter
eess.AS2024
LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
Wenhao Guan, Kaidi Wang, Wangjin Zhou +6
Recently, the application of diffusion models has facilitated the significant development of speech and audio generation. Nevertheless, the quality of samples generated by diffusio…
eess.AS2024★ 1 cited
MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score Prediction
Wangjin Zhou, Zhengdong Yang, Chenhui Chu +4
Automatic Mean Opinion Score (MOS) prediction is employed to evaluate the quality of synthetic speech. This study extends the application of predicted MOS to the task of Fake Audio…
eess.AS2023
LE-SSL-MOS: Self-Supervised Learning MOS Prediction with Listener Enhancement
Zili Qi, Xinhui Hu, Wangjin Zhou +4
Recently, researchers have shown an increasing interest in automatically predicting the subjective evaluation for speech synthesis systems. This prediction is a challenging task, e…