activity
20202026
most citedEnd-to-End Automatic Speech Recognition with Deep Mutual Learning

2 citations · 4 across the 19 of their papers we have counts for

collaborators
Showing cs.CLShow all

13 papers · 1 filter

cs.CL20231 cited

Transfer Learning from Pre-trained Language Models Improves End-to-End Speech Summarization

Kohei Matsuura, Takanori Ashihara, Takafumi Moriya +4

End-to-end speech summarization (E2E SSum) directly summarizes input speech into easy-to-read short sentences with a single model. This approach is promising because it, in contras…

cs.CL2023

End-to-End Joint Target and Non-Target Speakers ASR

Ryo Masumura, Naoki Makishima, Taiga Yamane +12

This paper proposes a novel automatic speech recognition (ASR) system that can transcribe individual speaker's speech while identifying whether they are target or non-target speake…

cs.CL2023

SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?

Takanori Ashihara, Takafumi Moriya, Kohei Matsuura +5

Self-supervised learning (SSL) for speech representation has been successfully applied in various downstream tasks, such as speech and speaker recognition. More recently, speech SS…

cs.CL2023

Exploration of Language Dependency for Japanese Self-Supervised Speech Representation Models

Takanori Ashihara, Takafumi Moriya, Kohei Matsuura +1

Self-supervised learning (SSL) has been dramatically successful not only in monolingual but also in cross-lingual settings. However, since the two settings have been studied indivi…

cs.CL20211 cited

End-to-End Rich Transcription-Style Automatic Speech Recognition with Semi-Supervised Learning

Tomohiro Tanaka, Ryo Masumura, Mana Ihori +3

We propose a semi-supervised learning method for building end-to-end rich transcription-style automatic speech recognition (RT-ASR) systems from small-scale rich transcription-styl…

cs.CL2021

Cross-Modal Transformer-Based Neural Correction Models for Automatic Speech Recognition

Tomohiro Tanaka, Ryo Masumura, Mana Ihori +5

We propose a cross-modal transformer-based neural correction models that refines the output of an automatic speech recognition (ASR) system so as to exclude ASR errors. Generally,…