28 citations · 44 across the 25 of their papers we have counts for
7 papers · 2 filters
Guided Speaker Embedding
Shota Horiguchi, Takafumi Moriya, Atsushi Ando +4
This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speaker…
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
Takafumi Moriya, Takanori Ashihara, Masato Mimura +4
A hybrid autoregressive transducer (HAT) is a variant of neural transducer that models blank and non-blank posterior distributions separately. In this paper, we propose a novel int…
Alignment-Free Training for Transducer-based Multi-Talker ASR
Takafumi Moriya, Shota Horiguchi, Marc Delcroix +5
Extending the RNN Transducer (RNNT) to recognize multi-talker speech is essential for wider automatic speech recognition (ASR) applications. Multi-talker RNNT (MT-RNNT) aims to ach…
NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando +15
We present a distant automatic speech recognition (DASR) system developed for the CHiME-8 DASR track. It consists of a diarization first pipeline. For diarization, we use end-to-en…
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
Shota Horiguchi, Atsushi Ando, Takafumi Moriya +4
This paper proposes a method for extracting speaker embedding for each speaker from a variable-length recording containing multiple speakers. Speaker embeddings are crucial not onl…
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
Hiroshi Sato, Takafumi Moriya, Masato Mimura +6
Real-time target speaker extraction (TSE) is intended to extract the desired speaker's voice from the observed mixture of multiple speakers in a streaming manner. Implementing real…