15 citations · 33 across the 20 of their papers we have counts for
28 papers
CasNet: Investigating Channel Robustness for Speech Separation
Fan-Lin Wang, Yao-Fei Cheng, Hung-Shin Lee +2
Recording channel mismatch between training and testing conditions has been shown to be a serious problem for speech separation. This situation greatly reduces the separation perfo…
Mandarin Singing Voice Synthesis with Denoising Diffusion Probabilistic Wasserstein GAN
Yin-Ping Cho, Yu Tsao, Hsin-Min Wang +1
Singing voice synthesis (SVS) is the computer production of a human-like singing voice from given musical scores. To accomplish end-to-end SVS effectively and efficiently, this wor…
Filter-based Discriminative Autoencoders for Children Speech Recognition
Chiang-Lin Tai, Hung-Shin Lee, Yu Tsao +1
Children speech recognition is indispensable but challenging due to the diversity of children's speech. In this paper, we propose a filter-based discriminative autoencoder for acou…
Generation of Speaker Representations Using Heterogeneous Training Batch Assembly
Yu-Huai Peng, Hung-Shin Lee, Pin-Tuan Huang +1
In traditional speaker diarization systems, a well-trained speaker model is a key component to extract representations from consecutive and partially overlapping segments in a long…
Subspace-based Representation and Learning for Phonotactic Spoken Language Recognition
Hung-Shin Lee, Yu Tsao, Shyh-Kang Jeng +1
Phonotactic constraints can be employed to distinguish languages by representing a speech utterance as a multinomial distribution or phone events. In the present study, we propose…
Partially Fake Audio Detection by Self-attention-based Fake Span Discovery
Haibin Wu, Heng-Cheng Kuo, Naijun Zheng +5
The past few years have witnessed the significant advances of speech synthesis and voice conversion technologies. However, such technologies can undermine the robustness of broadly…