15 citations · 27 across the 11 of their papers we have counts for
19 papers
Endpoint Detection for Streaming End-to-End Multi-talker ASR
Liang Lu, Jinyu Li, Yifan Gong
Streaming end-to-end multi-talker speech recognition aims at transcribing the overlapped speech from conversations or meetings with an all-neural model in a streaming fashion, whic…
Minimum Word Error Rate Training with Language Model Fusion for End-to-End Speech Recognition
Zhong Meng, Yu Wu, Naoyuki Kanda +6
Integrating external language models (LMs) into end-to-end (E2E) models remains a challenging task for domain-adaptive speech recognition. Recently, internal language model estimat…
Streaming Multi-talker Speech Recognition with Joint Speaker Identification
Liang Lu, Naoyuki Kanda, Jinyu Li +1
In multi-talker scenarios such as meetings and conversations, speech processing systems are usually required to transcribe the audio as well as identify the speakers for downstream…
Internal Language Model Training for Domain-Adaptive End-to-End Speech Recognition
Zhong Meng, Naoyuki Kanda, Yashesh Gaur +6
The efficacy of external language model (LM) integration with existing end-to-end (E2E) automatic speech recognition (ASR) systems can be improved significantly using the internal…
Minimum Bayes Risk Training for End-to-End Speaker-Attributed ASR
Naoyuki Kanda, Zhong Meng, Liang Lu +4
Recently, an end-to-end speaker-attributed automatic speech recognition (E2E SA-ASR) model was proposed as a joint model of speaker counting, speech recognition and speaker identif…
Internal Language Model Estimation for Domain-Adaptive End-to-End Speech Recognition
Zhong Meng, Sarangarajan Parthasarathy, Eric Sun +7
The external language models (LM) integration remains a challenging task for end-to-end (E2E) automatic speech recognition (ASR) which has no clear division between acoustic and la…