19 citations · 93 across the 56 of their papers we have counts for
7 papers · 2 filters
Mandarin Singing Voice Synthesis with Denoising Diffusion Probabilistic Wasserstein GAN
Yin-Ping Cho, Yu Tsao, Hsin-Min Wang +1
Singing voice synthesis (SVS) is the computer production of a human-like singing voice from given musical scores. To accomplish end-to-end SVS effectively and efficiently, this wor…
NASTAR: Noise Adaptive Speech Enhancement with Target-Conditional Resampling
Chi-Chang Lee, Cheng-Hung Hu, Yu-Chen Lin +3
For deep learning-based speech enhancement (SE) systems, the training-test acoustic mismatch can cause notable performance degradation. To address the mismatch issue, numerous nois…
A Study of Using Cepstrogram for Countermeasure Against Replay Attacks
Shih-Kuang Lee, Yu Tsao, Hsin-Min Wang
This study investigated the cepstrogram properties and demonstrated their effectiveness as powerful countermeasures against replay attacks. A cepstrum analysis of replay attacks su…
MTI-Net: A Multi-Target Speech Intelligibility Prediction Model
Ryandhimas E. Zezario, Szu-wei Fu, Fei Chen +3
Recently, deep learning (DL)-based non-intrusive speech assessment models have attracted great attention. Many studies report that these DL-based models yield satisfactory assessme…
MBI-Net: A Non-Intrusive Multi-Branched Speech Intelligibility Prediction Model for Hearing Aids
Ryandhimas E. Zezario, Fei Chen, Chiou-Shann Fuh +2
Improving the user's hearing ability to understand speech in noisy environments is critical to the development of hearing aid (HA) devices. For this, it is important to derive a me…
Partially Fake Audio Detection by Self-attention-based Fake Span Discovery
Haibin Wu, Heng-Cheng Kuo, Naijun Zheng +5
The past few years have witnessed the significant advances of speech synthesis and voice conversion technologies. However, such technologies can undermine the robustness of broadly…