activity
20182021
most citedConditional Teacher-Student Learning

108 citations · 290 across the 24 of their papers we have counts for

collaborators

32 papers

eess.AS2021

Continuous Speech Separation with Recurrent Selective Attention Network

Yixuan Zhang, Zhuo Chen, Jian Wu +4

While permutation invariant training (PIT) based continuous speech separation (CSS) significantly improves the conversation transcription accuracy, it often suffers from speech lea…

cs.CL20211 cited

Factorized Neural Transducer for Efficient Language Model Adaptation

Xie Chen, Zhong Meng, Sarangarajan Parthasarathy +1

In recent years, end-to-end (E2E) based automatic speech recognition (ASR) systems have achieved great success due to their simplicity and promising performance. Neural Transducer…

eess.AS20211 cited

A Comparative Study of Modular and Joint Approaches for Speaker-Attributed ASR on Monaural Long-Form Audio

Naoyuki Kanda, Xiong Xiao, Jian Wu +6

Speaker-attributed automatic speech recognition (SA-ASR) is a task to recognize "who spoke what" from multi-talker recordings. An SA-ASR system usually consists of multiple modules…

eess.AS2021

Minimum Word Error Rate Training with Language Model Fusion for End-to-End Speech Recognition

Zhong Meng, Yu Wu, Naoyuki Kanda +6

Integrating external language models (LMs) into end-to-end (E2E) models remains a challenging task for domain-adaptive speech recognition. Recently, internal language model estimat…

eess.AS20213 cited

Large-Scale Pre-Training of End-to-End Multi-Talker ASR for Meeting Transcription with Single Distant Microphone

Naoyuki Kanda, Guoli Ye, Yu Wu +5

Transcribing meetings containing overlapped speech with only a single distant microphone (SDM) has been one of the most challenging problems for automatic speech recognition (ASR).…

eess.AS20212 cited

End-to-End Speaker-Attributed ASR with Transformer

Naoyuki Kanda, Guoli Ye, Yashesh Gaur +4

This paper presents our recent effort on end-to-end speaker-attributed automatic speech recognition, which jointly performs speaker counting, speech recognition and speaker identif…