61 citations · 188 across the 42 of their papers we have counts for
9 papers · 1 filter
Investigation of Speaker Representation for Target-Speaker Speech Processing
Takanori Ashihara, Takafumi Moriya, Shota Horiguchi +5
Target-speaker speech processing (TS) tasks, such as target-speaker automatic speech recognition (TS-ASR), target speech extraction (TSE), and personal voice activity detection (p-…
Mamba-based Segmentation Model for Speaker Diarization
Alexis Plaquet, Naohiro Tawara, Marc Delcroix +3
Mamba is a newly proposed architecture which behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization,…
Guided Speaker Embedding
Shota Horiguchi, Takafumi Moriya, Atsushi Ando +4
This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speaker…
Alignment-Free Training for Transducer-based Multi-Talker ASR
Takafumi Moriya, Shota Horiguchi, Marc Delcroix +5
Extending the RNN Transducer (RNNT) to recognize multi-talker speech is essential for wider automatic speech recognition (ASR) applications. Multi-talker RNNT (MT-RNNT) aims to ach…
NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando +15
We present a distant automatic speech recognition (DASR) system developed for the CHiME-8 DASR track. It consists of a diarization first pipeline. For diarization, we use end-to-en…
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
Shota Horiguchi, Atsushi Ando, Takafumi Moriya +4
This paper proposes a method for extracting speaker embedding for each speaker from a variable-length recording containing multiple speakers. Speaker embeddings are crucial not onl…