3 papers
eess.AS2025
VBx for End-to-End Neural and Clustering-based Diarization
Petr Pálka, Jiangyu Han, Marc Delcroix +2
We present improvements to speaker diarization in the two-stage end-to-end neural diarization with vector clustering (EEND-VC) framework. The first stage employs a Conformer-based…
eess.AS2025
Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization
Jiangyu Han, Ruoyu Wang, Yoshiki Masuyama +4
Self-supervised models such as WavLM have demonstrated strong performance for neural speaker diarization. However, these models are typically pre-trained on single-channel recordin…
eess.AS2024
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
Alexander Polok, Dominik Klement, Martin Kocour +7
Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a significant challenge, particularly when systems conditioned on speaker embeddings fai…