activity
20172024
most citedLarge-Scale Pre-Training of End-to-End Multi-Talker ASR for Meeting Transcription with Single Distant Microphone

3 citations · 7 across the 6 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2024

Efficient Long-Form Speech Recognition for General Speech In-Context Learning

Hao Yen, Shaoshi Ling, Guoli Ye

We propose a novel approach to end-to-end automatic speech recognition (ASR) to achieve efficient speech in-context learning (SICL) for (i) long-form speech decoding, (ii) test-tim…

eess.AS20221 cited

Acoustic-aware Non-autoregressive Spell Correction with Mask Sample Decoding

Ruchao Fan, Guoli Ye, Yashesh Gaur +1

Masked language model (MLM) has been widely used for understanding tasks, e.g. BERT. Recently, MLM has also been used for generation tasks. The most popular one in speech is using…

eess.AS2021

Minimum Word Error Rate Training with Language Model Fusion for End-to-End Speech Recognition

Zhong Meng, Yu Wu, Naoyuki Kanda +6

Integrating external language models (LMs) into end-to-end (E2E) models remains a challenging task for domain-adaptive speech recognition. Recently, internal language model estimat…

eess.AS20213 cited

Large-Scale Pre-Training of End-to-End Multi-Talker ASR for Meeting Transcription with Single Distant Microphone

Naoyuki Kanda, Guoli Ye, Yu Wu +5

Transcribing meetings containing overlapped speech with only a single distant microphone (SDM) has been one of the most challenging problems for automatic speech recognition (ASR).…

eess.AS20212 cited

End-to-End Speaker-Attributed ASR with Transformer

Naoyuki Kanda, Guoli Ye, Yashesh Gaur +4

This paper presents our recent effort on end-to-end speaker-attributed automatic speech recognition, which jointly performs speaker counting, speech recognition and speaker identif…

eess.AS2020

Low Latency End-to-End Streaming Speech Recognition with a Scout Network

Chengyi Wang, Yu Wu, Shujie Liu +4

The attention-based Transformer model has achieved promising results for speech recognition (SR) in the offline mode. However, in the streaming mode, the Transformer model usually…