3 papers
eess.AS2025
VBx for End-to-End Neural and Clustering-based Diarization
Petr Pálka, Jiangyu Han, Marc Delcroix +2
We present improvements to speaker diarization in the two-stage end-to-end neural diarization with vector clustering (EEND-VC) framework. The first stage employs a Conformer-based…
eess.AS2025
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
Jiangyu Han, Petr Pálka, Marc Delcroix +4
Self-supervised learning (SSL) models such as WavLM have substantially advanced speaker diarization by providing rich contextual speech representations. However, the high computati…
eess.AS2024
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
Petr Pálka, Federico Landini, Dominik Klement +4
In spite of the popularity of end-to-end diarization systems nowadays, modular systems comprised of voice activity detection (VAD), speaker embedding extraction plus clustering, an…