4 citations · 6 across the 17 of their papers we have counts for
21 papers
Progressive Alignment Objectives for Aligner-Encoder based ASR
Jaeyoung Lee, Masato Mimura, Takafumi Moriya
Aligner-Encoders are recently proposed seq2seq end-to-end ASR models that replace decoder attention by predicting the uth token directly from the u-th encoder position, so the enco…
Chunkwise Aligners for Streaming Speech Recognition
Wen Shen Teo, Takafumi Moriya, Masato Mimura
We propose the Chunkwise Aligner, a novel architecture for streaming automatic speech recognition (ASR). While the Transducer is the standard model for streaming ASR, its training…
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
Takafumi Moriya, Masato Mimura, Tomohiro Tanaka +3
This paper proposes a unified framework, All-in-One ASR, that allows a single model to support multiple automatic speech recognition (ASR) paradigms, including connectionist tempor…
Generic Speech Enhancement with Self-Supervised Representation Space Loss
Hiroshi Sato, Tsubasa Ochiai, Marc Delcroix +3
Single-channel speech enhancement is utilized in various tasks to mitigate the effect of interfering signals. Conventionally, to ensure the speech enhancement performs optimally, t…
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando +15
In this paper, we introduce a multi-talker distant automatic speech recognition (DASR) system we designed for the DASR task 1 of the CHiME-8 challenge. Our system performs speaker…
Investigation of Speaker Representation for Target-Speaker Speech Processing
Takanori Ashihara, Takafumi Moriya, Shota Horiguchi +5
Target-speaker speech processing (TS) tasks, such as target-speaker automatic speech recognition (TS-ASR), target speech extraction (TSE), and personal voice activity detection (p-…