most citedLeveraging Large Text Corpora for End-to-End Speech Summarization

2 citations · 4 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2023

End-to-End Joint Target and Non-Target Speakers ASR

Ryo Masumura, Naoki Makishima, Taiga Yamane +12

This paper proposes a novel automatic speech recognition (ASR) system that can transcribe individual speaker's speech while identifying whether they are target or non-target speake…

eess.AS20231 cited

Knowledge Distillation for Neural Transducer-based Target-Speaker ASR: Exploiting Parallel Mixture/Single-Talker Speech Data

Takafumi Moriya, Hiroshi Sato, Tsubasa Ochiai +7

Neural transducer (RNNT)-based target-speaker speech recognition (TS-RNNT) directly transcribes a target speaker's voice from a multi-talker mixture. It is a promising approach for…

eess.AS2023

Improving Scheduled Sampling for Neural Transducer-based ASR

Takafumi Moriya, Takanori Ashihara, Hiroshi Sato +3

The recurrent neural network-transducer (RNNT) is a promising approach for automatic speech recognition (ASR) with the introduction of a prediction network that autoregressively co…

eess.AS2023

Downstream Task Agnostic Speech Enhancement with Self-Supervised Representation Loss

Hiroshi Sato, Ryo Masumura, Tsubasa Ochiai +8

Self-supervised learning (SSL) is the latest breakthrough in speech processing, especially for label-scarce downstream tasks by leveraging massive unlabeled audio data. The noise r…

cs.CL20232 cited

Leveraging Large Text Corpora for End-to-End Speech Summarization

Kohei Matsuura, Takanori Ashihara, Takafumi Moriya +4

End-to-end speech summarization (E2E SSum) is a technique to directly generate summary sentences from speech. Compared with the cascade approach, which combines automatic speech re…

cs.SD20221 cited

Speaker consistency loss and step-wise optimization for semi-supervised joint training of TTS and ASR using unpaired text data

Naoki Makishima, Satoshi Suzuki, Atsushi Ando +1

In this paper, we investigate the semi-supervised joint training of text to speech (TTS) and automatic speech recognition (ASR), where a small amount of paired data and a large amo…