3 citations · 11 across the 12 of their papers we have counts for
26 papers
Towards Lifelong Learning of End-to-end ASR
Heng-Jui Chang, Hung-yi Lee, Lin-shan Lee
Automatic speech recognition (ASR) technologies today are primarily optimized for given datasets; thus, any changes in the application environment (e.g., acoustic conditions or top…
End-to-end Whispered Speech Recognition with Frequency-weighted Approaches and Pseudo Whisper Pre-training
Heng-Jui Chang, Alexander H. Liu, Hung-yi Lee +1
Whispering is an important mode of human speech, but no end-to-end recognition results for it were reported yet, probably due to the scarcity of available whispered speech data. In…
Interrupted and cascaded permutation invariant training for speech separation
Gene-Ping Yang, Szu-Lin Wu, Yao-Wen Mao +2
Permutation Invariant Training (PIT) has long been a stepping stone method for training speech separation model in handling the label ambiguity problem. With PIT selecting the mini…
Sequence-to-sequence Automatic Speech Recognition with Word Embedding Regularization and Fused Decoding
Alexander H. Liu, Tzu-Wei Sung, Shun-Po Chuang +2
In this paper, we investigate the benefit that off-the-shelf word embedding can bring to the sequence-to-sequence (seq-to-seq) automatic speech recognition (ASR). We first introduc…
Towards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation Learning
Alexander H. Liu, Tao Tu, Hung-yi Lee +1
In this paper we propose a Sequential Representation Quantization AutoEncoder (SeqRQ-AE) to learn from primarily unpaired audio data and produce sequences of representations very c…
Improved Speech Separation with Time-and-Frequency Cross-domain Joint Embedding and Clustering
Gene-Ping Yang, Chao-I Tuan, Hung-Yi Lee +1
Speech separation has been very successful with deep learning techniques. Substantial effort has been reported based on approaches over spectrogram, which is well known as the stan…