activity
20202022
most citedMandarin-English Code-switching Speech Recognition with Self-supervised Speech Representation Models

8 citations · 12 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CL20221 cited

SpeechCLIP: Integrating Speech with Pre-Trained Vision and Language Model

Yi-Jen Shih, Hsuan-Fu Wang, Heng-Jui Chang +3

Data-driven speech processing models usually perform well with a large amount of text supervision, but collecting transcribed speech data is costly. Therefore, we propose SpeechCLI…

cs.CL20223 cited

SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities

Hsiang-Sheng Tsai, Heng-Jui Chang, Wen-Chin Huang +14

Transfer learning has proven to be crucial in advancing the state of speech and natural language processing research in recent years. In speech, a model pre-trained by self-supervi…

cs.CL20218 cited

Mandarin-English Code-switching Speech Recognition with Self-supervised Speech Representation Models

Liang-Hsuan Tseng, Yu-Kuan Fu, Heng-Jui Chang +1

Code-switching (CS) is common in daily conversations where more than one language is used within a sentence. The difficulties of CS speech recognition lie in alternating languages…

cs.CL2021

Non-autoregressive Mandarin-English Code-switching Speech Recognition

Shun-Po Chuang, Heng-Jui Chang, Sung-Feng Huang +1

Mandarin-English code-switching (CS) is frequently used among East and Southeast Asian people. However, the intra-sentence language switching of the two very different languages ma…

cs.CL2021

Towards Lifelong Learning of End-to-end ASR

Heng-Jui Chang, Hung-yi Lee, Lin-shan Lee

Automatic speech recognition (ASR) technologies today are primarily optimized for given datasets; thus, any changes in the application environment (e.g., acoustic conditions or top…

cs.CL2020

End-to-end Whispered Speech Recognition with Frequency-weighted Approaches and Pseudo Whisper Pre-training

Heng-Jui Chang, Alexander H. Liu, Hung-yi Lee +1

Whispering is an important mode of human speech, but no end-to-end recognition results for it were reported yet, probably due to the scarcity of available whispered speech data. In…