activity
20192026
most citedJointly optimal denoising, dereverberation, and source separation

45 citations · 129 across the 53 of their papers we have counts for

collaborators
Showing eess.ASShow all

52 papers · 1 filter

eess.AS2025

Reference Microphone Selection for Guided Source Separation based on the Normalized L-p Norm

Anselm Lohmann, Tomohiro Nakatani, Rintaro Ikeshita +3

Guided Source Separation (GSS) is a popular front-end for distant automatic speech recognition (ASR) systems using spatially distributed microphones. When considering spatially dis…

eess.AS2025

Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?

Shota Horiguchi, Naohiro Tawara, Takanori Ashihara +2

Neural speaker diarization is widely used for overlap-aware speaker diarization, but it requires large multi-speaker datasets for training. To meet this data requirement, large dat…

eess.AS2025

MOVER: Combining Multiple Meeting Recognition Systems

Naoyuki Kamo, Tsubasa Ochiai, Marc Delcroix +1

In this paper, we propose Meeting recognizer Output Voting Error Reduction (MOVER), a novel system combination method for meeting recognition tasks. Although there are methods to c…

eess.AS20254 cited

Generic Speech Enhancement with Self-Supervised Representation Space Loss

Hiroshi Sato, Tsubasa Ochiai, Marc Delcroix +3

Single-channel speech enhancement is utilized in various tasks to mitigate the effect of interfering signals. Conventionally, to ensure the speech enhancement performs optimally, t…

eess.AS2025

Mitigating Non-Target Speaker Bias in Guided Speaker Embedding

Shota Horiguchi, Takanori Ashihara, Marc Delcroix +2

Obtaining high-quality speaker embeddings in multi-speaker conditions is crucial for many applications. A recently proposed guided speaker embedding framework, which utilizes speec…

eess.AS2025

Pretraining Multi-Speaker Identification for Neural Speaker Diarization

Shota Horiguchi, Atsushi Ando, Marc Delcroix +1

End-to-end speaker diarization enables accurate overlap-aware diarization by jointly estimating multiple speakers' speech activities in parallel. This approach is data-hungry, requ…