activity
20182022
most citedMetricGAN: Generative Adversarial Networks based Black-box Metric Scores Optimization for Speech Enhancement

117 citations · 233 across the 37 of their papers we have counts for

collaborators

65 papers

cs.CL2022

T5lephone: Bridging Speech and Text Self-supervised Models for Spoken Language Understanding via Phoneme level T5

Chan-Jan Hsu, Ho-Lam Chung, Hung-yi Lee +1

In Spoken language understanding (SLU), a natural solution is concatenating pre-trained speech models (e.g. HuBERT) and pretrained language models (PLM, e.g. T5). Most previous wor…

cs.SD20221 cited

CasNet: Investigating Channel Robustness for Speech Separation

Fan-Lin Wang, Yao-Fei Cheng, Hung-Shin Lee +2

Recording channel mismatch between training and testing conditions has been shown to be a serious problem for speech separation. This situation greatly reduces the separation perfo…

eess.AS20221 cited

Mandarin Singing Voice Synthesis with Denoising Diffusion Probabilistic Wasserstein GAN

Yin-Ping Cho, Yu Tsao, Hsin-Min Wang +1

Singing voice synthesis (SVS) is the computer production of a human-like singing voice from given musical scores. To accomplish end-to-end SVS effectively and efficiently, this wor…

cs.CL20223 cited

Filter-based Discriminative Autoencoders for Children Speech Recognition

Chiang-Lin Tai, Hung-Shin Lee, Yu Tsao +1

Children speech recognition is indispensable but challenging due to the diversity of children's speech. In this paper, we propose a filter-based discriminative autoencoder for acou…

cs.CL2022

XDBERT: Distilling Visual Information to BERT from Cross-Modal Systems to Improve Language Understanding

Chan-Jan Hsu, Hung-yi Lee, Yu Tsao

Transformer-based models are widely used in natural language understanding (NLU) tasks, and multimodal transformers have been effective in visual-language tasks. This study explore…

eess.AS20222 cited

Partial Coupling of Optimal Transport for Spoken Language Identification

Xugang Lu, Peng Shen, Yu Tsao +1

In order to reduce domain discrepancy to improve the performance of cross-domain spoken language identification (SLID) system, as an unsupervised domain adaptation (UDA) method, we…