most citedLeveraging Large Text Corpora for End-to-End Speech Summarization

2 citations · 5 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CL20231 cited

Transfer Learning from Pre-trained Language Models Improves End-to-End Speech Summarization

Kohei Matsuura, Takanori Ashihara, Takafumi Moriya +4

End-to-end speech summarization (E2E SSum) directly summarizes input speech into easy-to-read short sentences with a single model. This approach is promising because it, in contras…

eess.AS20231 cited

Knowledge Distillation for Neural Transducer-based Target-Speaker ASR: Exploiting Parallel Mixture/Single-Talker Speech Data

Takafumi Moriya, Hiroshi Sato, Tsubasa Ochiai +7

Neural transducer (RNNT)-based target-speaker speech recognition (TS-RNNT) directly transcribes a target speaker's voice from a multi-talker mixture. It is a promising approach for…

eess.AS2023

Improving Scheduled Sampling for Neural Transducer-based ASR

Takafumi Moriya, Takanori Ashihara, Hiroshi Sato +3

The recurrent neural network-transducer (RNNT) is a promising approach for automatic speech recognition (ASR) with the introduction of a prediction network that autoregressively co…

eess.AS2023

Downstream Task Agnostic Speech Enhancement with Self-Supervised Representation Loss

Hiroshi Sato, Ryo Masumura, Tsubasa Ochiai +8

Self-supervised learning (SSL) is the latest breakthrough in speech processing, especially for label-scarce downstream tasks by leveraging massive unlabeled audio data. The noise r…

cs.CL2023

Exploration of Language Dependency for Japanese Self-Supervised Speech Representation Models

Takanori Ashihara, Takafumi Moriya, Kohei Matsuura +1

Self-supervised learning (SSL) has been dramatically successful not only in monolingual but also in cross-lingual settings. However, since the two settings have been studied indivi…

cs.CL20232 cited

Leveraging Large Text Corpora for End-to-End Speech Summarization

Kohei Matsuura, Takanori Ashihara, Takafumi Moriya +4

End-to-end speech summarization (E2E SSum) is a technique to directly generate summary sentences from speech. Compared with the cascade approach, which combines automatic speech re…