activity
20172026
most citedEncoder-Decoder Based Attractors for End-to-End Neural Diarization

61 citations · 188 across the 42 of their papers we have counts for

collaborators
Showing 2024Show all

9 papers · 1 filter

cs.SD2024

Investigation of Speaker Representation for Target-Speaker Speech Processing

Takanori Ashihara, Takafumi Moriya, Shota Horiguchi +5

Target-speaker speech processing (TS) tasks, such as target-speaker automatic speech recognition (TS-ASR), target speech extraction (TSE), and personal voice activity detection (p-…

cs.SD2024

Mamba-based Segmentation Model for Speaker Diarization

Alexis Plaquet, Naohiro Tawara, Marc Delcroix +3

Mamba is a newly proposed architecture which behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization,…

eess.AS2024★ 1 cited

Guided Speaker Embedding

Shota Horiguchi, Takafumi Moriya, Atsushi Ando +4

This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speaker…

eess.AS2024

Alignment-Free Training for Transducer-based Multi-Talker ASR

Takafumi Moriya, Shota Horiguchi, Marc Delcroix +5

Extending the RNN Transducer (RNNT) to recognize multi-talker speech is essential for wider automatic speech recognition (ASR) applications. Multi-talker RNNT (MT-RNNT) aims to ach…

eess.AS2024

NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge

Naoyuki Kamo, Naohiro Tawara, Atsushi Ando +15

We present a distant automatic speech recognition (DASR) system developed for the CHiME-8 DASR track. It consists of a diarization first pipeline. For diarization, we use end-to-en…

eess.AS2024

Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings

Shota Horiguchi, Atsushi Ando, Takafumi Moriya +4

This paper proposes a method for extracting speaker embedding for each speaker from a variable-length recording containing multiple speakers. Speaker embeddings are crucial not onl…