collaborators

10 papers

eess.AS2025

Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models

Jiangyu Han, Petr Pálka, Marc Delcroix +4

Self-supervised learning (SSL) models such as WavLM have substantially advanced speaker diarization by providing rich contextual speech representations. However, the high computati…

eess.AS2025

Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization

Jiangyu Han, Federico Landini, Johan Rohdin +4

Self-supervised learning (SSL) models like WavLM can be effectively utilized when building speaker diarization systems but are often large and slow, limiting their use in resource…

eess.AS2025

Analysis of ABC Frontend Audio Systems for the NIST-SRE24

Sara Barahona, Anna Silnova, Ladislav Mošner +14

We present a comprehensive analysis of the embedding extractors (frontends) developed by the ABC team for the audio track of NIST SRE 2024. We follow the two scenarios imposed by N…

eess.AS2024

DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition

Alexander Polok, Dominik Klement, Martin Kocour +7

Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a significant challenge, particularly when systems conditioned on speaker embeddings fai…

eess.AS2024

Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization

Petr Pálka, Federico Landini, Dominik Klement +4

In spite of the popularity of end-to-end diarization systems nowadays, modular systems comprised of voice activity detection (VAD), speaker embedding extraction plus clustering, an…

eess.AS2024

Leveraging Self-Supervised Learning for Speaker Diarization

Jiangyu Han, Federico Landini, Johan Rohdin +3

End-to-end neural diarization has evolved considerably over the past few years, but data scarcity is still a major obstacle for further improvements. Self-supervised learning metho…