activity
20202022
most citedSurpriseNet: Melody Harmonization Conditioning on User-controlled Surprise Contours

5 citations · 9 across the 7 of their papers we have counts for

collaborators

8 papers

cs.SD20221 cited

CasNet: Investigating Channel Robustness for Speech Separation

Fan-Lin Wang, Yao-Fei Cheng, Hung-Shin Lee +2

Recording channel mismatch between training and testing conditions has been shown to be a serious problem for speech separation. This situation greatly reduces the separation perfo…

cs.CL20223 cited

Filter-based Discriminative Autoencoders for Children Speech Recognition

Chiang-Lin Tai, Hung-Shin Lee, Yu Tsao +1

Children speech recognition is indispensable but challenging due to the diversity of children's speech. In this paper, we propose a filter-based discriminative autoencoder for acou…

cs.SD2022

Generation of Speaker Representations Using Heterogeneous Training Batch Assembly

Yu-Huai Peng, Hung-Shin Lee, Pin-Tuan Huang +1

In traditional speaker diarization systems, a well-trained speaker model is a key component to extract representations from consecutive and partially overlapping segments in a long…

cs.SD2022

Subspace-based Representation and Learning for Phonotactic Spoken Language Recognition

Hung-Shin Lee, Yu Tsao, Shyh-Kang Jeng +1

Phonotactic constraints can be employed to distinguish languages by representing a speech utterance as a multinomial distribution or phone events. In the present study, we propose…

cs.SD20215 cited

SurpriseNet: Melody Harmonization Conditioning on User-controlled Surprise Contours

Yi-Wei Chen, Hung-Shin Lee, Yen-Hsing Chen +1

The surprisingness of a song is an essential and seemingly subjective factor in determining whether the listener likes it. With the help of information theory, it can be described…

eess.AS2021

Relational Data Selection for Data Augmentation of Speaker-dependent Multi-band MelGAN Vocoder

Yi-Chiao Wu, Cheng-Hung Hu, Hung-Shin Lee +5

Nowadays, neural vocoders can generate very high-fidelity speech when a bunch of training data is available. Although a speaker-dependent (SD) vocoder usually outperforms a speaker…