activity
20202026
most citedLeveraging Large Text Corpora for End-to-End Speech Summarization

2 citations · 8 across the 18 of their papers we have counts for

collaborators

19 papers

eess.AS2026

Improving Large-Scale Weakly Supervised ASR by Filtering and Selection

Kohei Matsuura, Masato Mimura

Leveraging large-scale weakly supervised datasets is crucial to train robust end-to-end automatic speech recognition (ASR) models. However, such datasets often contain noisy labels…

cs.SD2026

Frontend Token Enhancement for Token-Based Speech Recognition

Takanori Ashihara, Shota Horiguchi, Kohei Matsuura +2

Discretized representations of speech signals are efficient alternatives to continuous features for various speech applications, including automatic speech recognition (ASR) and sp…

eess.AS2025

Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge

Naoyuki Kamo, Naohiro Tawara, Atsushi Ando +15

In this paper, we introduce a multi-talker distant automatic speech recognition (DASR) system we designed for the DASR task 1 of the CHiME-8 challenge. Our system performs speaker…

cs.SD2024

Investigation of Speaker Representation for Target-Speaker Speech Processing

Takanori Ashihara, Takafumi Moriya, Shota Horiguchi +5

Target-speaker speech processing (TS) tasks, such as target-speaker automatic speech recognition (TS-ASR), target speech extraction (TSE), and personal voice activity detection (p-…

eess.AS2024

Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding

Takafumi Moriya, Takanori Ashihara, Masato Mimura +4

A hybrid autoregressive transducer (HAT) is a variant of neural transducer that models blank and non-blank posterior distributions separately. In this paper, we propose a novel int…

eess.AS2024

Alignment-Free Training for Transducer-based Multi-Talker ASR

Takafumi Moriya, Shota Horiguchi, Marc Delcroix +5

Extending the RNN Transducer (RNNT) to recognize multi-talker speech is essential for wider automatic speech recognition (ASR) applications. Multi-talker RNNT (MT-RNNT) aims to ach…