most citedSpeaker Adaptation for Attention-Based End-to-End Speech Recognition

41 citations · 77 across the 8 of their papers we have counts for

collaborators

8 papers

eess.AS20203 cited

Investigation of End-To-End Speaker-Attributed ASR for Continuous Multi-Talker Recordings

Naoyuki Kanda, Xuankai Chang, Yashesh Gaur +4

Recently, an end-to-end (E2E) speaker-attributed automatic speech recognition (SA-ASR) model was proposed as a joint model of speaker counting, speech recognition and speaker ident…

cs.LG20208 cited

Federated Transfer Learning with Dynamic Gradient Aggregation

Dimitrios Dimitriadis, Kenichi Kumatani, Robert Gmyr +2

In this paper, a Federated Learning (FL) simulation platform is introduced. The target scenario is Acoustic Model training based on this platform. To our knowledge, this is the fir…

eess.AS20205 cited

Joint Speaker Counting, Speech Recognition, and Speaker Identification for Overlapped Speech of Any Number of Speakers

Naoyuki Kanda, Yashesh Gaur, Xiaofei Wang +4

We propose an end-to-end speaker-attributed automatic speech recognition model that unifies speaker counting, speech recognition, and speaker identification on monaural overlapped…

cs.CL20201 cited

Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR

Hirofumi Inaguma, Yashesh Gaur, Liang Lu +2

Recently, a few novel streaming attention-based sequence-to-sequence (S2S) models have been proposed to perform online speech recognition with linear-time decoding complexity. Howe…

eess.AS202018 cited

On the Comparison of Popular End-to-End Models for Large Scale Speech Recognition

Jinyu Li, Yu Wu, Yashesh Gaur +3

Recently, there has been a strong push to transition from hybrid models to end-to-end (E2E) models for automatic speech recognition. Currently, there are three promising E2E method…

eess.AS20201 cited

Domain Adaptation via Teacher-Student Learning for End-to-End Speech Recognition

Zhong Meng, Jinyu Li, Yashesh Gaur +1

Teacher-student (T/S) has shown to be effective for domain adaptation of deep neural network acoustic models in hybrid speech recognition systems. In this work, we extend the T/S l…