most citedKnowledge Distillation for Neural Transducer-based Target-Speaker ASR: Exploiting Parallel Mixture/Single-Talker Speech Data

1 citations · 1 across the 6 of their papers we have counts for

collaborators

6 papers

eess.AS2024

Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding

Takafumi Moriya, Takanori Ashihara, Masato Mimura +4

A hybrid autoregressive transducer (HAT) is a variant of neural transducer that models blank and non-blank posterior distributions separately. In this paper, we propose a novel int…

eess.AS2024

Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings

Shota Horiguchi, Atsushi Ando, Takafumi Moriya +4

This paper proposes a method for extracting speaker embedding for each speaker from a variable-length recording containing multiple speakers. Speaker embeddings are crucial not onl…

cs.CL2023

End-to-End Joint Target and Non-Target Speakers ASR

Ryo Masumura, Naoki Makishima, Taiga Yamane +12

This paper proposes a novel automatic speech recognition (ASR) system that can transcribe individual speaker's speech while identifying whether they are target or non-target speake…

eess.AS20231 cited

Knowledge Distillation for Neural Transducer-based Target-Speaker ASR: Exploiting Parallel Mixture/Single-Talker Speech Data

Takafumi Moriya, Hiroshi Sato, Tsubasa Ochiai +7

Neural transducer (RNNT)-based target-speaker speech recognition (TS-RNNT) directly transcribes a target speaker's voice from a multi-talker mixture. It is a promising approach for…

eess.AS2023

Improving Scheduled Sampling for Neural Transducer-based ASR

Takafumi Moriya, Takanori Ashihara, Hiroshi Sato +3

The recurrent neural network-transducer (RNNT) is a promising approach for automatic speech recognition (ASR) with the introduction of a prediction network that autoregressively co…

eess.AS2023

Downstream Task Agnostic Speech Enhancement with Self-Supervised Representation Loss

Hiroshi Sato, Ryo Masumura, Tsubasa Ochiai +8

Self-supervised learning (SSL) is the latest breakthrough in speech processing, especially for label-scarce downstream tasks by leveraging massive unlabeled audio data. The noise r…