3 citations · 7 across the 6 of their papers we have counts for
6 papers · 1 filter
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
Hao Yen, Shaoshi Ling, Guoli Ye
We propose a novel approach to end-to-end automatic speech recognition (ASR) to achieve efficient speech in-context learning (SICL) for (i) long-form speech decoding, (ii) test-tim…
Acoustic-aware Non-autoregressive Spell Correction with Mask Sample Decoding
Ruchao Fan, Guoli Ye, Yashesh Gaur +1
Masked language model (MLM) has been widely used for understanding tasks, e.g. BERT. Recently, MLM has also been used for generation tasks. The most popular one in speech is using…
Minimum Word Error Rate Training with Language Model Fusion for End-to-End Speech Recognition
Zhong Meng, Yu Wu, Naoyuki Kanda +6
Integrating external language models (LMs) into end-to-end (E2E) models remains a challenging task for domain-adaptive speech recognition. Recently, internal language model estimat…
Large-Scale Pre-Training of End-to-End Multi-Talker ASR for Meeting Transcription with Single Distant Microphone
Naoyuki Kanda, Guoli Ye, Yu Wu +5
Transcribing meetings containing overlapped speech with only a single distant microphone (SDM) has been one of the most challenging problems for automatic speech recognition (ASR).…
End-to-End Speaker-Attributed ASR with Transformer
Naoyuki Kanda, Guoli Ye, Yashesh Gaur +4
This paper presents our recent effort on end-to-end speaker-attributed automatic speech recognition, which jointly performs speaker counting, speech recognition and speaker identif…
Low Latency End-to-End Streaming Speech Recognition with a Scout Network
Chengyi Wang, Yu Wu, Shujie Liu +4
The attention-based Transformer model has achieved promising results for speech recognition (SR) in the offline mode. However, in the streaming mode, the Transformer model usually…