collaborators

6 papers

eess.AS2026

Improving Large-Scale Weakly Supervised ASR by Filtering and Selection

Kohei Matsuura, Masato Mimura

Leveraging large-scale weakly supervised datasets is crucial to train robust end-to-end automatic speech recognition (ASR) models. However, such datasets often contain noisy labels…

eess.AS2026

Progressive Alignment Objectives for Aligner-Encoder based ASR

Jaeyoung Lee, Masato Mimura, Takafumi Moriya

Aligner-Encoders are recently proposed seq2seq end-to-end ASR models that replace decoder attention by predicting the uth token directly from the u-th encoder position, so the enco…

eess.AS2026

Chunkwise Aligners for Streaming Speech Recognition

Wen Shen Teo, Takafumi Moriya, Masato Mimura

We propose the Chunkwise Aligner, a novel architecture for streaming automatic speech recognition (ASR). While the Transducer is the standard model for streaming ASR, its training…

eess.AS2026

Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR

Jaeyoung Lee, Masato Mimura

We present a decoder-only Conformer for automatic speech recognition (ASR) that processes speech and text in a single stack without external speech encoders or pretrained large lan…

eess.AS2025

All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR

Takafumi Moriya, Masato Mimura, Tomohiro Tanaka +3

This paper proposes a unified framework, All-in-One ASR, that allows a single model to support multiple automatic speech recognition (ASR) paradigms, including connectionist tempor…

eess.AS2025

Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge

Naoyuki Kamo, Naohiro Tawara, Atsushi Ando +15

In this paper, we introduce a multi-talker distant automatic speech recognition (DASR) system we designed for the DASR task 1 of the CHiME-8 challenge. Our system performs speaker…