most citedSpeech-language Pre-training for End-to-end Spoken Language Understanding

9 citations · 25 across the 5 of their papers we have counts for

collaborators

6 papers

cs.SD20222 cited

Deploying self-supervised learning in the wild for hybrid automatic speech recognition

Mostafa Karimi, Changliang Liu, Kenichi Kumatani +3

Self-supervised learning (SSL) methods have proven to be very successful in automatic speech recognition (ASR). These great improvements have been reported mostly based on highly c…

cs.SD20211 cited

Improving Noise Robustness of Contrastive Speech Representation Learning with Speech Reconstruction

Heming Wang, Yao Qian, Xiaofei Wang +6

Noise robustness is essential for deploying automatic speech recognition (ASR) systems in real-world environments. One way to reduce the effect of noise interference is to employ a…

cs.CL20214 cited

Multilingual Speech Recognition using Knowledge Transfer across Learning Processes

Rimita Lahiri, Kenichi Kumatani, Eric Sun +1

Multilingual end-to-end(E2E) models have shown a great potential in the expansion of the language coverage in the realm of automatic speech recognition(ASR). In this paper, we aim…

eess.AS20219 cited

UniSpeech at scale: An Empirical Study of Pre-training Method on Large-Scale Speech Recognition Dataset

Chengyi Wang, Yu Wu, Shujie Liu +4

Recently, there has been a vast interest in self-supervised learning (SSL) where the model is pre-trained on large scale unlabeled data and then fine-tuned on a small labeled datas…

cs.CL20219 cited

Speech-language Pre-training for End-to-end Spoken Language Understanding

Yao Qian, Ximo Bian, Yu Shi +4

End-to-end (E2E) spoken language understanding (SLU) can infer semantics directly from speech signal without cascading an automatic speech recognizer (ASR) with a natural language…

cs.CL2021

UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data

Chengyi Wang, Yu Wu, Yao Qian +5

In this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both unlabeled and labeled data, in which supervised phonetic CTC le…