117 citations · 233 across the 37 of their papers we have counts for
65 papers
T5lephone: Bridging Speech and Text Self-supervised Models for Spoken Language Understanding via Phoneme level T5
Chan-Jan Hsu, Ho-Lam Chung, Hung-yi Lee +1
In Spoken language understanding (SLU), a natural solution is concatenating pre-trained speech models (e.g. HuBERT) and pretrained language models (PLM, e.g. T5). Most previous wor…
CasNet: Investigating Channel Robustness for Speech Separation
Fan-Lin Wang, Yao-Fei Cheng, Hung-Shin Lee +2
Recording channel mismatch between training and testing conditions has been shown to be a serious problem for speech separation. This situation greatly reduces the separation perfo…
Mandarin Singing Voice Synthesis with Denoising Diffusion Probabilistic Wasserstein GAN
Yin-Ping Cho, Yu Tsao, Hsin-Min Wang +1
Singing voice synthesis (SVS) is the computer production of a human-like singing voice from given musical scores. To accomplish end-to-end SVS effectively and efficiently, this wor…
Filter-based Discriminative Autoencoders for Children Speech Recognition
Chiang-Lin Tai, Hung-Shin Lee, Yu Tsao +1
Children speech recognition is indispensable but challenging due to the diversity of children's speech. In this paper, we propose a filter-based discriminative autoencoder for acou…
XDBERT: Distilling Visual Information to BERT from Cross-Modal Systems to Improve Language Understanding
Chan-Jan Hsu, Hung-yi Lee, Yu Tsao
Transformer-based models are widely used in natural language understanding (NLU) tasks, and multimodal transformers have been effective in visual-language tasks. This study explore…
Partial Coupling of Optimal Transport for Spoken Language Identification
Xugang Lu, Peng Shen, Yu Tsao +1
In order to reduce domain discrepancy to improve the performance of cross-domain spoken language identification (SLID) system, as an unsupervised domain adaptation (UDA) method, we…