activity
20172025
most citedSpeaker Adaptation for Attention-Based End-to-End Speech Recognition

41 citations · 102 across the 23 of their papers we have counts for

collaborators
Showing eess.ASShow all

15 papers · 1 filter

eess.AS2024

Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens

Jinzheng Zhao, Niko Moritz, Egor Lakomkin +7

Cascaded speech-to-speech translation systems often suffer from the error accumulation problem and high latency, which is a result of cascaded modules whose inference delays accumu…

eess.AS20221 cited

Acoustic-aware Non-autoregressive Spell Correction with Mask Sample Decoding

Ruchao Fan, Guoli Ye, Yashesh Gaur +1

Masked language model (MLM) has been widely used for understanding tasks, e.g. BERT. Recently, MLM has also been used for generation tasks. The most popular one in speech is using…

eess.AS20211 cited

A Comparative Study of Modular and Joint Approaches for Speaker-Attributed ASR on Monaural Long-Form Audio

Naoyuki Kanda, Xiong Xiao, Jian Wu +6

Speaker-attributed automatic speech recognition (SA-ASR) is a task to recognize "who spoke what" from multi-talker recordings. An SA-ASR system usually consists of multiple modules…

eess.AS20213 cited

Large-Scale Pre-Training of End-to-End Multi-Talker ASR for Meeting Transcription with Single Distant Microphone

Naoyuki Kanda, Guoli Ye, Yu Wu +5

Transcribing meetings containing overlapped speech with only a single distant microphone (SDM) has been one of the most challenging problems for automatic speech recognition (ASR).…

eess.AS20212 cited

End-to-End Speaker-Attributed ASR with Transformer

Naoyuki Kanda, Guoli Ye, Yashesh Gaur +4

This paper presents our recent effort on end-to-end speaker-attributed automatic speech recognition, which jointly performs speaker counting, speech recognition and speaker identif…

eess.AS20211 cited

Internal Language Model Training for Domain-Adaptive End-to-End Speech Recognition

Zhong Meng, Naoyuki Kanda, Yashesh Gaur +6

The efficacy of external language model (LM) integration with existing end-to-end (E2E) automatic speech recognition (ASR) systems can be improved significantly using the internal…